
Senior Software Engineer, NeMo Core Platform
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in California, +5 more states.
• Develop the NeMo Platform, NVIDIA’s solution for creating, assessing, deploying, and managing AI systems on a large scale.
• Design an Agentic Execution framework for local environments, Kubernetes, Slurm, on-premises setups, and air-gapped contexts.
• Provide senior technical guidance through design evaluations, code assessments, mentorship, and ownership of complex cross-component challenges.
• Construct and uphold Core Platform APIs for job execution, data storage, entity management, secrets handling, RBAC, and authentication.
• Enhance the plugin architecture to enable teams and external clients to integrate new functionalities.
• Contribute to open development in NVIDIA’s open-source repository.
• Deploy production-ready code utilizing agentic coding tools.
• Enhance the reliability, observability, debuggability, and performance across the platform, SDKs, plugins, jobs, and developer workflows.
• Establish robust unit, integration, end-to-end, Docker, and Kubernetes test coverage.
• BS, MS, or equivalent experience in Computer Science, Computer Engineering, or a related technical domain.
• Over 10 years of professional software engineering experience in building production systems.
• Comfort in navigating a fast-paced and ambiguous environment.
• Outstanding verbal and written communication abilities.
• Capability to produce and evaluate high-quality architectural RFCs.
• Proficient system design skills.
• Strong grasp of reliability, scalability, security, and performance trade-offs in production infrastructure.
• Experience with distributed systems, cloud-native services, containers, Kubernetes, and job orchestration.
• Excellent Python programming skills, including API design, typing, testing, debugging, performance analysis, and maintainable software design.
• Experience in designing SDKs, libraries, plugins, CLIs, or other developer-facing interfaces.
• Ability to work autonomously, define technical scope, break down ambiguous issues, and drive work across team boundaries.
• Experience in building, deploying, and iterating on production agentic AI systems at scale in Kubernetes.
• Familiarity with advanced plugin architectures.
• Ability to link technical evaluation efforts to business outcomes, product quality, user experience, reliability, or operational efficiency.
• Experience with enterprise AI systems that involve measurement, regression testing, observability, governance, and ongoing improvement for production deployment.
• Equity
• Benefits
opinov8
IGS Energy
Cast & Crew
FCamara Consulting & Training
Get handpicked remote jobs straight to your inbox weekly.