
AI Cluster Architect
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Design and architect extensive GPU clusters while adhering to fixed site power budgets to optimize GPU density, leaving room for compute services, storage, and networking.
• Model and assess power consumption across GPUs, CPUs, NICs, switches, fabric components, storage, and facility constraints.
• Analyze InfiniBand, RoCE, and SpectrumX networking architectures, including multi-plane, 2-tier/3-tier, and rail-optimized topologies.
• Identify network scale limitations based on switch radix, link speed, topology, and blocking specifications.
• Collect, analyze, and maintain SKU-level power and thermal data for GPUs, NICs, switches, DPUs, storage, and server platforms.
• Create power-aware cluster configuration templates and capacity-planning models tailored for sites with diverse constraints.
• Document the architecture, design decisions, trade-off analyses, and operational considerations for deployment and lifecycle management.
• Offer insights on the integration of next-generation GPUs, NICs, and fabrics.
• Work in partnership with vendors on innovative fabric architectures for large-scale cluster deployments exceeding 100k GPUs.
• A minimum of 7 years of experience in designing or constructing large-scale HPC, AI, or hyperscale GPU clusters.
• In-depth knowledge of GPU and accelerator system design, including node topology, PCIe, NVLink, NVSwitch, ROCm, and NIC-to-GPU affinity.
• Strong expertise in InfiniBand, RoCE, and SpectrumX networking.
• Experience with multi-tier, multi-plane, Clos/dragonfly variants, and large-radix switch design.
• Proven track record in modeling power draw and thermal characteristics of servers, GPUs, NICs, switches, optics, and storage.
• Capability to design fully non-blocking networks or intentionally manage over-/under-subscription while understanding the impact on workload performance.
• Demonstrated ability to gather and analyze vendor SKU-level specifications and integrate them into scalable cluster architectures.
• Experience in balancing customer-driven compute, storage, and service density requirements with the total GPU count.
• Strong skills in documentation, communication, and cross-functional collaboration.
• Must reside in a U.S. state and be legally authorized to work in the United States.
• Must indicate whether employment visa sponsorship is currently or potentially needed.
• Comprehensive Medical Benefits with 100% company-paid premiums for employee-only plans.
• Fully covered dental and vision premiums at no cost to employees.
• 401(k) plan with a 100% match up to 4%, featuring immediate vesting.
• Annual Professional Development Reimbursement of $2,500.
• 11 recognized Holidays.
• Accrual of Paid Time Off (PTO).
• PTO Rollover Plan available.
• Enjoy your birthday off as a holiday.
• Increased PTO granted at 3-year and 10-year anniversaries.
• One month of paid sabbatical every five years.
• Anniversary Bonus provided each year.
• $500 allocated for first-year remote office setup.
• $400 each subsequent year for new equipment.
• Internet reimbursement up to $75 monthly.
• Gym membership reimbursement of up to $50 monthly.
• Company-sponsored Wellable subscription.
Ensemble Health Partners
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.