
Senior MLOps Engineer – DSX Enablement
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in Poland, +2 more countries.
• Create pioneering solutions that enhance AI infrastructure capabilities.
• Provide guidance to infrastructure specialists on the requirements of ML workloads.
• Assist practitioners in identifying and resolving issues within full-stack AI and ML systems.
• Support both internal and external clients' AI and ML projects, including evaluating LLM performance and integrating new hardware within open-source frameworks.
• Design and implement tailored AI solutions on NeoCloud platforms and through NVIDIA Cloud Partners, which include distributed training, inference optimization, and MLOps pipelines.
• Serve as the primary technical liaison for customers and partners, facilitate collaborative efforts, ensure the success of initiatives on DGX Cloud, and address complex production challenges.
• Collaborate closely with infrastructure software and accelerated-framework teams.
• Analyze and optimize large-scale training and inference workloads to minimize latency, costs, and operational risks.
• Create open-source tools and reference architectures for machine learning and AI workloads, pipelines, and systems at scale.
• Bachelor’s, Master’s, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical discipline, or equivalent professional experience.
• Over 8 years of experience in technical positions such as data science, data engineering, or ML engineering, preferably focusing on large-scale production environments.
• Proven experience in AI/ML across various stages of the machine learning lifecycle, from exploratory analysis to production system deployment.
• Proficiency with Linux, batch schedulers, Kubernetes, distributed filesystems, and advanced networking at a datacenter level.
• Strong scripting and programming abilities in Bash and Python.
• Solid systems programming expertise in C++, Go, or Rust.
• Experience utilizing machine learning or deep learning frameworks for both training and inference tasks.
• Exceptional communication and technical presentation abilities, with a talent for conveying architectures, trade-offs, and recommendations to both engineering and leadership audiences.
• A demonstrated track record of engineering discipline and successful project execution.
• Experience contributing to and engaging with open-source communities.
• Familiarity with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, RAPIDS, Triton, NIM, InfiniBand, NVLink, and RoCE.
• Experience in developing machine learning systems in security-sensitive environments, alongside distributed training and inference frameworks.
• Understanding of cloud-native MLOps methodologies, encompassing containerization, CI/CD, workflow automation, observability stacks, and GitOps practices.
• Direct experience in diagnosing and resolving cross-layer performance or correctness issues that involve hardware, networking, accelerators, hypervisors or operating systems, compilers or runtimes, application code, and libraries.
• Competitive salaries.
• Generous benefits package.
Stack AV
ICA, Inc.
Reddit, Inc.
Upstart
Get handpicked remote jobs straight to your inbox weekly.