
Datacenter Infrastructure Specialist
Posted Aug 21

Posted Aug 21
This is a fully remote position, open to applicants in United States.
• Validate new hardware to ensure that partner deployments align with Runpod specifications for distributed AI/ML workloads.
• Monitor the health of the fleet and pinpoint any performance decline.
• Audit downtime and deliver technical data to safeguard customer SLAs.
• Leverage LLMs and AI agents to automate network triage and create dynamic fleet runbooks.
• Manage technical incident communications and convert outages into actionable solutions.
• Facilitate the growth of infrastructure partners.
• Own the technical lifecycle and operational well-being of Runpod’s high-density GPU fleet.
• Advise, translate for, onboard, and serve as incident commander for hardware partners.
• Connect hardware partners with internal engineering teams.
• Contribute to the resilience, scalability, and revenue growth of Runpod’s global physical infrastructure.
• 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
• Strong expertise in standard datacenter networking and performance troubleshooting.
• Familiarity with RDMA, InfiniBand, or RoCE is highly preferred.
• Practical experience with the NVIDIA Software Stack, including driver installation and performance utilities.
• Understanding of multi-node performance tuning.
• Strong Linux system administration skills.
• Experience with containerization technologies, including Docker.
• System-level troubleshooting and performance tuning at the kernel and hardware interface layers.
• Excellent written and verbal communication abilities.
• Willingness to participate in a future on-call rotation.
• Detail-oriented and proactive in identifying potential failures.
• Preferred: experience in startups, HPC bare-metal environments at massive scale, Grafana, Prometheus, Datadog, Python, Go, Bash, and internal APIs.
• Must be eligible to work in the United States; Runpod is currently unable to sponsor employment visas.
• Meaningful equity; all team members receive stock options.
• Comprehensive medical, dental & vision plans; 100% coverage for employees and partial coverage for dependents.
• Flexible PTO.
• Remote-first work environment fostering inclusive, collaborative teams.
• $1,200 Home Office & Equipment Stipend.
• Opportunities for culture, learning, and ownership.
Hitachi Solutions America
Nava
OpenTeams
Kalepa
Get handpicked remote jobs straight to your inbox weekly.