
Senior Manager, AI Infrastructure Operations
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Lead the engineering team tasked with the daily execution, scaling, and management of AI compute clusters.
• Transform engineering roadmaps and technical specifications from the Director of AI Infrastructure into comprehensive project plans and execution benchmarks.
• Drive the deployment of clusters, hardware initialization, node setup, and integration with orchestration and scheduling systems.
• Ensure the reliability, availability, and performance of clusters through effective monitoring, automation, and ongoing operational enhancements.
• Oversee lifecycle management for bare metal and GPU fleets, which includes provisioning, configuration management, firmware/driver updates, and hardware validation.
• Manage incident response for GPU and cluster infrastructure, guaranteeing prompt resolution and thorough root-cause analysis.
• Collaborate closely with AI/ML, SRE, Networking, and Hardware Engineering teams to ensure cluster capabilities fulfill training and inference requirements.
• Coordinate with Product teams to verify technical specifications, feature readiness, and delivery schedules.
• Support integrations across networking, storage, scheduling, and resource orchestration components.
• Enhance tools and automation for cluster provisioning, observability, configuration management, and large-scale fleet operations.
• Contribute to the design and improvement of multi-tenant scheduling, workload management, and orchestration systems in collaboration with senior technical personnel.
• Identify performance limitations and suggest engineering-level optimizations.
• Mentor and guide engineers, cultivating a high-performance, detail-oriented engineering culture.
• Assist with career growth, expectations, and performance evaluations for team members.
• Help refine engineering practices, including code reviews, testing standards, documentation, and operational runbooks.
• 6–10 years of experience in infrastructure engineering, HPC, large-scale systems, or related fields.
• Strong understanding of AI compute infrastructure, including GPU/CPU clusters, distributed training architectures, and high-performance networking (InfiniBand/RDMA).
• Experience managing production bare metal, GPU, or hardware fleet operations at significant scale.
• Practical expertise with Linux systems, Kubernetes or Slurm, provisioning tools (Terraform, Ansible), observability platforms, and networking fundamentals.
• Proven experience in cluster operations, hardware initialization, distributed systems, or ML workload support.
• Experience leading engineering teams or groups, with the capability to oversee execution while remaining engaged with technical tasks.
• Ability to effectively communicate with cross-functional engineering teams and convert strategy into actionable engineering tasks.
• Strong execution-oriented mindset with the ability to prioritize, deliver, and adapt in a dynamic environment.
• Excellent Medical Benefits with 100% company-paid premiums for employee-only plans, plus 100% company-paid dental and vision premiums.
• 401(k) plan with a 100% match up to 4% and immediate vesting.
• Professional Development Reimbursement of $2,500 each year.
• 11 Holidays plus Paid Time Off Accrual, Rollover Plan, and the option to take your birthday off.
• Increased PTO at 3-year and 10-year anniversaries, along with a 1-month paid sabbatical every 5 years and an Anniversary Bonus each year.
• $500 for first-year remote office setup and $400 each subsequent year for new equipment.
• Internet reimbursement of up to $75 per month.
• Gym membership reimbursement of up to $50 per month.
• Company-paid Wellable subscription.
Grupo Boticário
Asurion
Tiger Analytics
knowmad mood
Get handpicked remote jobs straight to your inbox weekly.