Datacenter Infrastructure Specialist

atRunPodRemoteUS flagUnited StatesFull-timeInfrastructure EngineerMid-levelSenior$120k – $160k/year

Posted Aug 21

This is a fully remote position, open to applicants in United States.

📋 Description

• Validate new hardware to ensure that partner deployments align with Runpod specifications for distributed AI/ML workloads.

• Monitor the health of the fleet and pinpoint any performance decline.

• Audit downtime and deliver technical data to safeguard customer SLAs.

• Leverage LLMs and AI agents to automate network triage and create dynamic fleet runbooks.

• Manage technical incident communications and convert outages into actionable solutions.

• Facilitate the growth of infrastructure partners.

• Own the technical lifecycle and operational well-being of Runpod’s high-density GPU fleet.

• Advise, translate for, onboard, and serve as incident commander for hardware partners.

• Connect hardware partners with internal engineering teams.

• Contribute to the resilience, scalability, and revenue growth of Runpod’s global physical infrastructure.


⛳️ Requirements

• 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.

• Strong expertise in standard datacenter networking and performance troubleshooting.

• Familiarity with RDMA, InfiniBand, or RoCE is highly preferred.

• Practical experience with the NVIDIA Software Stack, including driver installation and performance utilities.

• Understanding of multi-node performance tuning.

• Strong Linux system administration skills.

• Experience with containerization technologies, including Docker.

• System-level troubleshooting and performance tuning at the kernel and hardware interface layers.

• Excellent written and verbal communication abilities.

• Willingness to participate in a future on-call rotation.

• Detail-oriented and proactive in identifying potential failures.

• Preferred: experience in startups, HPC bare-metal environments at massive scale, Grafana, Prometheus, Datadog, Python, Go, Bash, and internal APIs.

• Must be eligible to work in the United States; Runpod is currently unable to sponsor employment visas.


🏝️ Benefits

• Meaningful equity; all team members receive stock options.

• Comprehensive medical, dental & vision plans; 100% coverage for employees and partial coverage for dependents.

• Flexible PTO.

• Remote-first work environment fostering inclusive, collaborative teams.

• $1,200 Home Office & Equipment Stipend.

• Opportunities for culture, learning, and ownership.

People also viewed

Hitachi Solutions America23 hours ago

Senior Azure Infrastructure Architect

DE flagGermany, +1 more countryFull-timeInfrastructure Engineer$140k – $180k/year
ApplyView job
Nava1 day ago

Senior Infrastructure Engineer, Azure

US flagAlabama, +29 more statesFull-timeInfrastructure Engineer$135.9k – $153k/year
ApplyView job
OpenTeams1 day ago

Senior Infrastructure Engineer – AI/ML Platform

US flagUnited States OnlyFull-timeInfrastructure Engineer$145k – $250k/year
ApplyView job
Kalepa1 day ago

Staff Platform Engineer – Infrastructure

US flagUnited States OnlyFull-timeInfrastructure Engineer$220k – $300k/year
ApplyView job
ElevenLabs1 day ago

HPC Infrastructure Engineer – GPU Clusters

US flagUnited States OnlyFull-timeInfrastructure Engineer
ApplyView job
Activision Blizzard1 day ago

Senior Azure Infrastructure Engineer

US flagCalifornia OnlyFull-timeInfrastructure Engineer$102.8k – $190.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers