
Senior Engineer
Posted Jul 23

Posted Jul 23
This is a fully remote position, open to applicants in California, +1 more state.
β’ Create innovative solutions that enhance AI infrastructure capabilities.
β’ Have a direct impact on customer success through pioneering AI initiatives.
β’ Design and implement custom AI solutions on NCP and Neo Cloud platforms, which include distributed training, inference optimization, and MLOps pipelines built on NVIDIA reference architectures.
β’ Serve as the primary technical liaison for strategic NCPs, providing both remote and on-site support, resolving complex production challenges, and advising partner engineering teams on NVIDIA platform protocols.
β’ Deploy and manage AI workloads across DGX Cloud, NCP data centers, and leading CSP environments using Kubernetes, containers, and GPU scheduling systems tailored to NCP builds.
β’ Profile and optimize large-scale training and inference workloads on NCP platforms.
β’ Implement observability and SLO/SLA monitoring.
β’ Lead comprehensive initiatives to minimize latency, cost, and operational risk.
β’ Expand and implement NVIDIA reference architectures on partner platforms, establish integrations with partner control planes and customer environments, and ensure seamless API, data pipeline, and enterprise software connectivity.
β’ Develop detailed implementation guides, runbooks, and post-mortem documentation that encapsulate best practices for executing NVIDIA AI workloads at scale on NCP platforms.
β’ BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.
β’ Over 8 years of experience in customer-facing technical roles such as Solutions Engineering, DevOps, Site Reliability, or ML Infrastructure Engineering, preferably in large-scale cloud or service provider environments.
β’ Strong proficiency in Linux systems, distributed computing, Kubernetes, containers, and GPU scheduling on multi-tenant or service provider platforms.
β’ Proven AI/ML experience supporting large-scale training and inference workloads (e.g., LLMs, generative models, recommendation systems) in production or mission-critical settings.
β’ Solid programming abilities in Python/Go, with practical experience using frameworks such as PyTorch or TensorFlow for training and deployment.
β’ Demonstrated ability to work collaboratively with customer and partner engineering teams in dynamic environments, guiding complex technical investigations and resolving issues to their root causes.
β’ Exceptional communication and technical presentation skills, with the capability to clearly articulate architectures, trade-offs, and recommendations to both engineering and leadership audiences.
β’ Equity
β’ Benefits
TMS
StaffMill Oy
HumanIT Digital Consulting
HumanIT Digital Consulting
Get handpicked remote jobs straight to your inbox weekly.