
Datacenter Infrastructure Specialist
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in United States.
• Assist in validating new hardware to ensure that partner deployments align with Runpod’s specifications for distributed AI/ML workloads.
• Monitor fleet health to detect performance issues. You will audit downtime and provide the necessary technical data to uphold customer SLAs.
• We adopt an AI-first approach, leveraging the technology we host to enhance our operations. You will collaborate with LLMs and AI agents to automate network triage and create dynamic runbooks for our fleet.
• Facilitate technical incident communications with clear updates, serving as a reliable point of contact that translates outages into actionable solutions.
• Contribute to the growth of our infrastructure partners.
• 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.
• Strong expertise in standard datacenter networking and performance troubleshooting. Familiarity with RDMA, InfiniBand, or RoCE is highly preferred.
• Practical experience with the NVIDIA Software Stack (driver installation, performance utilities) and understanding of multi-node performance optimization.
• Proficient in Linux system administration and containerization (Docker), with comfort in performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers.
• Excellent written and verbal communication skills, capable of articulating hardware or networking issues to both technical partners and internal leadership.
• As our global fleet expands, you may be required to participate in an on-call rotation in the future.
• Detail-oriented and proactive in identifying potential failures before they affect customers.
• Experience working in a fast-paced environment and contributing to the development of operational workflows.
• Background in managing or optimizing bare-metal High-Performance Computing environments at large scale.
• Familiarity with Grafana, Prometheus, or Datadog for system health monitoring.
• Proficient in Python, Go (Golang), or Bash for automating repetitive infrastructure tasks and interfacing with internal APIs.
• Meaningful equity in a rapidly growing AI infrastructure company—everyone on the team receives stock options—your contributions drive our growth, and you share in the rewards.
• Generous medical, dental, and vision plans—100% coverage for all employees and partial coverage for dependents.
• Flexible PTO—take the time you need to recharge.
• Most roles are remote-first, with inclusive, collaborative teams using Slack as the primary mode of internal communication.
• Join a dedicated team at the forefront of AI infrastructure—where culture, learning, and ownership are central to our scaling efforts.
• $1,200 Home Office & Equipment Stipend—We equip you for success from day one with the gear and support needed to create your ideal workspace.
Teleperformance
Trilon Group
Carbon60
fal
Get handpicked remote jobs straight to your inbox weekly.