Remotery

Senior Manager, AI Infrastructure Operations

atVultrRemoteUS flagUnited StatesFull-timeLLM EngineerSenior$150k – $160k/year

Posted 4 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Lead the engineering team tasked with the daily execution, scaling, and management of AI compute clusters.

• Transform engineering roadmaps and technical specifications from the Director of AI Infrastructure into comprehensive project plans and execution benchmarks.

• Drive the deployment of clusters, hardware initialization, node setup, and integration with orchestration and scheduling systems.

• Ensure the reliability, availability, and performance of clusters through effective monitoring, automation, and ongoing operational enhancements.

• Oversee lifecycle management for bare metal and GPU fleets, which includes provisioning, configuration management, firmware/driver updates, and hardware validation.

• Manage incident response for GPU and cluster infrastructure, guaranteeing prompt resolution and thorough root-cause analysis.

• Collaborate closely with AI/ML, SRE, Networking, and Hardware Engineering teams to ensure cluster capabilities fulfill training and inference requirements.

• Coordinate with Product teams to verify technical specifications, feature readiness, and delivery schedules.

• Support integrations across networking, storage, scheduling, and resource orchestration components.

• Enhance tools and automation for cluster provisioning, observability, configuration management, and large-scale fleet operations.

• Contribute to the design and improvement of multi-tenant scheduling, workload management, and orchestration systems in collaboration with senior technical personnel.

• Identify performance limitations and suggest engineering-level optimizations.

• Mentor and guide engineers, cultivating a high-performance, detail-oriented engineering culture.

• Assist with career growth, expectations, and performance evaluations for team members.

• Help refine engineering practices, including code reviews, testing standards, documentation, and operational runbooks.


⛳️ Requirements

• 6–10 years of experience in infrastructure engineering, HPC, large-scale systems, or related fields.

• Strong understanding of AI compute infrastructure, including GPU/CPU clusters, distributed training architectures, and high-performance networking (InfiniBand/RDMA).

• Experience managing production bare metal, GPU, or hardware fleet operations at significant scale.

• Practical expertise with Linux systems, Kubernetes or Slurm, provisioning tools (Terraform, Ansible), observability platforms, and networking fundamentals.

• Proven experience in cluster operations, hardware initialization, distributed systems, or ML workload support.

• Experience leading engineering teams or groups, with the capability to oversee execution while remaining engaged with technical tasks.

• Ability to effectively communicate with cross-functional engineering teams and convert strategy into actionable engineering tasks.

• Strong execution-oriented mindset with the ability to prioritize, deliver, and adapt in a dynamic environment.


🏝️ Benefits

• Excellent Medical Benefits with 100% company-paid premiums for employee-only plans, plus 100% company-paid dental and vision premiums.

• 401(k) plan with a 100% match up to 4% and immediate vesting.

• Professional Development Reimbursement of $2,500 each year.

• 11 Holidays plus Paid Time Off Accrual, Rollover Plan, and the option to take your birthday off.

• Increased PTO at 3-year and 10-year anniversaries, along with a 1-month paid sabbatical every 5 years and an Anniversary Bonus each year.

• $500 for first-year remote office setup and $400 each subsequent year for new equipment.

• Internet reimbursement of up to $75 per month.

• Gym membership reimbursement of up to $50 per month.

• Company-paid Wellable subscription.

People also viewed

Grupo Boticário11 hours ago

Specialist I, Generative AI and Agents Engineer

Anywhere in the WorldFull-timeLLM Engineer
ApplyView job
Asurion11 hours ago

Staff Software Engineer – Generative AI Engineering

US flagTennessee, +1 more stateFull-timeLLM Engineer
ApplyView job
Tiger Analytics3 days ago

Forward Deployed Engineer, Generative AI

US flagTexas OnlyFull-timeLLM Engineer
ApplyView job
knowmad mood3 days ago

Senior Fullstack AI Engineer, LLM, RAG, Knowledge Graphs

ES flagSpain OnlyFull-timeLLM Engineer
ApplyView job
3Pillar Global4 days ago

AI Lead Engineer, Prompt Engineering, RAG, LLMs

IN flagIndia OnlyFreelanceLLM Engineer
ApplyView job
Extractta4 days ago

Engenheiro de IA Generativa Sênior

Anywhere in the WorldFull-timeLLM Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers