
Senior MLOps Engineer – DSX Enablement
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California, +1 more state.
• Create cutting-edge solutions to enhance AI infrastructure capabilities.
• Provide guidance to infrastructure specialists regarding the requirements of ML workloads.
• Assist practitioners in diagnosing and resolving full-stack AI and ML system issues.
• Support both internal and external clients’ AI and ML projects, which include evaluating LLM performance and integrating new hardware into open-source frameworks.
• Design and implement tailored AI solutions on NeoCloud platforms and with NVIDIA Cloud Partners, covering distributed training, inference optimization, and MLOps pipelines.
• Serve as the main technical liaison for internal and external clients and partners.
• Oversee collaborative efforts, ensure successful initiatives on DGX Cloud, and address complex production challenges.
• Work alongside infrastructure software and accelerated-framework teams.
• Analyze and optimize large-scale training and inference workloads on NVIDIA Cloud Partner platforms.
• Lead initiatives to minimize latency, costs, and operational risks.
• Create open-source tools and reference architectures for machine learning and AI workloads, pipelines, and systems at scale.
• Bachelor's, Master's, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical discipline, or equivalent experience.
• Over 8 years of experience in technical roles such as data science, data engineering, or ML engineering, ideally focused on large-scale production systems.
• Proven AI/ML experience through various stages of the machine learning lifecycle, from exploratory analysis to deployment.
• Proficient in Linux, batch schedulers, Kubernetes, distributed filesystems, and advanced networking at datacenter scale.
• Strong scripting and programming abilities in Bash and Python.
• Solid systems programming expertise in C++, Go, or Rust.
• Experience with machine learning or deep learning frameworks for both training and inference.
• Exceptional communication and technical presentation skills, capable of explaining architectures, trade-offs, and recommendations to engineering and leadership audiences.
• A demonstrable record of engineering discipline and successful execution on engaging projects.
• Experience in contributing to and collaborating within open-source communities.
• Familiarity with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, RAPIDS, Triton, NIM, InfiniBand, NVLink, and RoCE.
• Experience in developing machine learning systems within security-critical environments and utilizing distributed training and inference frameworks.
• Understanding of MLOps practices in a cloud-native environment, encompassing containerization, CI/CD pipelines, workflow automation, observability stacks, and GitOps workflows.
• Direct experience in diagnosing and resolving performance or correctness issues across hardware, networking, accelerators, hypervisors, operating systems, compilers, runtimes, application code, and libraries.
• Competitive salaries.
• Generous benefits package.
• Equity.
NVIDIA
SentiLink
SentiLink
Leega
Get handpicked remote jobs straight to your inbox weekly.