
Distinguished Engineer, Production Engineering, Data Center Automation
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California, +3 more states.
β’ Establish the long-term technical strategy for the consistent operation of DGX Cloud clusters across on-premises, hyperscalers, and NeoCloud environments.
β’ Define the architectural vision and essential operational guidelines for the lifecycle of clusters, runtime delivery, restoration, release readiness, and ongoing operability of DGX Cloud resources.
β’ Lead the roadmap and implementation of key cross-organizational investments that enhance production readiness, operational safety, performance, and collaboration across teams.
β’ Make and guide significant technical decisions regarding the coordination of platform, hardware, provider, and service teams in the operation of DGX Cloud resources in a production setting.
β’ Create robust workflows, interfaces, and foster engineering collaboration across Kubernetes production services, provider and hardware readiness, on-premises and bare-metal infrastructure operations, as well as service-layer reliability domains.
β’ Serve as a senior technical leader within the Production Engineering team.
β’ Establish the architectural vision for cluster operations in DGX Cloud.
β’ Set operational standards and oversee the evolution of the production model.
β’ Drive the delivery of cross-organizational capabilities to ensure DGX Cloud resources remain usable, maintainable, and continuously scalable.
β’ Lead by influence across multiple teams to achieve critical production outcomes.
β’ BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical discipline, or equivalent experience.
β’ Over 18 years of experience in building and managing large-scale distributed systems, infrastructure platforms, or production environments.
β’ Proven company-level technical leadership at a principal, distinguished, or equivalent level in production engineering, SRE, infrastructure software, or cloud platforms.
β’ Demonstrated experience in developing operating models, architectural direction, and engineering standards across diverse technical domains and organizations.
β’ A strong track record of leading extensive, cross-team technical initiatives from conception to production, including aligning collaborators, managing complexities, and delivering measurable results.
β’ In-depth expertise in software engineering, system knowledge, and production insights.
β’ Familiarity with Kubernetes service management.
β’ Experience with on-premises, hyperscaler, NeoCloud, and bare-metal infrastructure operations.
β’ Proven ability to develop automation, workflows, interfaces, APIs, architectures, or operating standards for large-scale infrastructure environments.
β’ Equity
β’ Benefits
EXL
Revecore
Capgemini
Get handpicked remote jobs straight to your inbox weekly.