Remotery

Datacenter Infrastructure Specialist

Posted Jul 29

This is a fully remote position, open to applicants in United States.

đź“‹ Description

• Assist in validating new hardware to ensure that partner deployments align with Runpod’s specifications for distributed AI/ML workloads.

• Monitor fleet health to detect performance issues. You will audit downtime and provide the necessary technical data to uphold customer SLAs.

• We adopt an AI-first approach, leveraging the technology we host to enhance our operations. You will collaborate with LLMs and AI agents to automate network triage and create dynamic runbooks for our fleet.

• Facilitate technical incident communications with clear updates, serving as a reliable point of contact that translates outages into actionable solutions.

• Contribute to the growth of our infrastructure partners.


⛳️ Requirements

• 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.

• Strong expertise in standard datacenter networking and performance troubleshooting. Familiarity with RDMA, InfiniBand, or RoCE is highly preferred.

• Practical experience with the NVIDIA Software Stack (driver installation, performance utilities) and understanding of multi-node performance optimization.

• Proficient in Linux system administration and containerization (Docker), with comfort in performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers.

• Excellent written and verbal communication skills, capable of articulating hardware or networking issues to both technical partners and internal leadership.

• As our global fleet expands, you may be required to participate in an on-call rotation in the future.

• Detail-oriented and proactive in identifying potential failures before they affect customers.

• Experience working in a fast-paced environment and contributing to the development of operational workflows.

• Background in managing or optimizing bare-metal High-Performance Computing environments at large scale.

• Familiarity with Grafana, Prometheus, or Datadog for system health monitoring.

• Proficient in Python, Go (Golang), or Bash for automating repetitive infrastructure tasks and interfacing with internal APIs.


🏝️ Benefits

• Meaningful equity in a rapidly growing AI infrastructure company—everyone on the team receives stock options—your contributions drive our growth, and you share in the rewards.

• Generous medical, dental, and vision plans—100% coverage for all employees and partial coverage for dependents.

• Flexible PTO—take the time you need to recharge.

• Most roles are remote-first, with inclusive, collaborative teams using Slack as the primary mode of internal communication.

• Join a dedicated team at the forefront of AI infrastructure—where culture, learning, and ownership are central to our scaling efforts.

• $1,200 Home Office & Equipment Stipend—We equip you for success from day one with the gear and support needed to create your ideal workspace.

People also viewed

Teleperformance14 hours ago

Senior Systems Engineer – IT Infrastructure Engineer

HR flagCroatia OnlyFull-timeInfrastructure Engineer
ApplyView job
Trilon Group1 day ago

AWS Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$90k – $110k/year
ApplyView job
Carbon601 day ago

Principal Infrastructure Architect

CA flagCanada OnlyFull-timeInfrastructure EngineerC$180k – C$220k/year
ApplyView job
fal1 day ago

Senior/Staff Kubernetes Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer$180k – $250k/year
ApplyView job
LatamCent1 day ago

Lead Security and Infrastructure Engineer

US flagFlorida OnlyFull-timeInfrastructure Engineer$170k – $210k/year
ApplyView job
AIP Publishing1 day ago

Cloud Infrastructure Engineer

US flagConnecticut, +9 more statesFull-timeInfrastructure Engineer$125k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers