
Principal Technologist – AI Compute
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the technical architecture for intricate, high-stakes customer projects involving AI/HPC GPU and accelerator compute cluster design, workload performance, and the compute-to-network interface.
• Collaborate with Sales and Solutions Architects from initial discovery to deal closure.
• Lead the design and execution of proof-of-concept projects.
• Establish criteria for proof-of-concept success and develop benchmarking plans for GPU clusters focused on training/inference throughput and scaling efficiency.
• Present proof-of-concept outcomes to customers with precision and clarity.
• Interact with ML infrastructure leads, compute architects, and executive stakeholders at hyperscalers, NeoClouds, service providers, and large enterprises.
• Gather insights from the field regarding GPU/accelerator platform trends and workload behavior.
• Convert field insights into actionable requirements for Product Management and Engineering.
• Create and distribute architecture playbooks, reference designs, and best practices for AI compute infrastructure.
• Mentor and guide Solutions Architects and Solutions Engineers.
• Represent DriveNets at industry conferences and events.
• Write white papers and technical blogs.
• Enhance DriveNets' external technical reputation in the AI compute infrastructure sector.
• Over 12 years of experience in designing and architecting data center compute infrastructure.
• A minimum of 3 years dedicated to AI/HPC GPU or accelerator platforms in hyperscale settings.
• Comprehensive hands-on experience with GPU/accelerator hardware and system design, particularly with NVIDIA/AMD.
• Proficient in compute cluster orchestration utilizing Kubernetes and Slurm.
• Familiarity with distributed training/inference frameworks.
• Demonstrated experience in senior pre-sales, solutions architecture, or system architecture roles.
• Direct involvement in influencing large, complex deals with VP and C-level executives.
• Knowledge of GPU virtualization and partitioning, including MIG/vGPU.
• Experience in managing containerized ML workloads and provisioning bare-metal GPUs.
• Skilled in scripting and automation with Python, APIs, and JSON.
• Outstanding communication and presentation abilities.
• Willingness to travel domestically and internationally up to 10% of the time.
• Familiar with PyTorch, TensorFlow, and JAX.
• Experience with NCCL/RCCL tuning and GPU cluster benchmarking, such as MLPerf.
• Understanding of collective communication dynamics at scale.
• Practical knowledge of scale-up interconnects, including NVLink and UALink.
• Familiarity with GPU resource scheduling and orchestration using Slurm, Kubernetes, and Ray.
• Experience in GPU observability and telemetry with tools like DCGM, Prometheus, and Grafana.
• Knowledge of data center operations fundamentals, including power, cooling, and rack design.
• Holds a BS, MS, PhD, or equivalent experience in Electrical/Computer Engineering, Computer Science, Physics, or another Engineering discipline.
• Opportunity to significantly impact multi-million-dollar customer projects.
• A creative, collaborative, and growth-focused work environment.
• Approximately 10% travel, both domestic and international.
• Commitment to equal opportunity employment.
• Advantageous NVIDIA/AMD platform certifications or equivalent qualifications.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.