
Technical Staff Member – Datacenter Operations
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in California.
• Take charge of the operational readiness of the physical infrastructure supporting the GPU cloud.
• Oversee rack deployment, cabling, inventory management, and acceptance testing for new GPU capacity in collaboration with datacenter partners and engineering teams.
• Keep asset records, rack layouts, power allocations, cabling documentation, and spare parts inventories up to date.
• Direct hardware fault triage and manage remote hands, vendor escalations, component replacements, and RMA processes.
• Develop maintenance plans and change procedures that minimize disruptions for customers while safeguarding equipment and data.
• Monitor capacity readiness, trends in hardware failures, repair durations, and operational risks.
• Streamline repetitive reporting and operational workflows through automation.
• Collaborate with facility teams on power, cooling, environmental monitoring, and preparation for high-density GPU deployments.
• Create runbooks and escalation procedures.
• Assist with incident response across datacenter and infrastructure teams.
• Coordinate deployments, hardware maintenance, and incident responses with datacenter partners.
• Transform new capacity into reliable production infrastructure and decrease repair times.
• A minimum of 3 years of experience in datacenter operations, hardware infrastructure, or production systems operations.
• Practical experience in deploying and troubleshooting rack-mounted servers, networking devices, and structured cabling.
• Proven experience in coordinating with datacenter providers, remote hands, and hardware vendors during deployments and incidents.
• Familiarity with Linux diagnostics, BMC consoles, and server hardware health monitoring tools.
• Exceptional operational judgment, strong documentation practices, and a commitment to resolving issues.
• Understanding of GPU server components, PCIe devices, memory, storage, and hardware diagnostics.
• Knowledgeable in rack power budgeting, redundant power paths, airflow, and principles of high-density cooling.
• Proficient in fiber and copper cabling, optics, labeling, and physical network troubleshooting.
• Experienced in asset tracking, spare parts management, change control, and incident management.
• Basic scripting skills for inventory management, health checks, and operational automation.
• Awareness of safe datacenter working practices.
• Experience with NVIDIA DGX/HGX systems or large GPU cluster deployments is advantageous.
• Familiarity with liquid-cooled infrastructure and coordination with facility engineering is a plus.
• Experience in establishing new datacenter sites or expanding multi-site capacity is a plus.
• Experience in hardware qualification, burn-in testing, and reliability analysis is an advantage.
• Knowledge of integrating physical operations with automated fleet provisioning is a plus.
• Equity incentives.
• Direct engagement with customers and a world-class engineering team.
• Opportunity to influence systems that drive next-generation AI breakthroughs.
Vitable Health
The Cigna Group
AmpiFire
Worldwide Clinical Trials
Get handpicked remote jobs straight to your inbox weekly.