
Senior AI Compute Engineer
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in California.
• Deploy, manage, and validate AI Compute/HPC infrastructure within Linux-based environments for both new and existing clients.
• Serve as the subject matter expert during planning discussions and through to implementation with customers.
• Execute documentation and knowledge transfer related to handovers to assist clients in deploying complex systems.
• Provide insights to internal teams by reporting bugs, documenting workarounds, and recommending enhancements.
• Collaborate with customers, partners, and internal teams to analyze, define, and implement large-scale AI Compute initiatives that encompass networking, system design, automation, and validation.
• Over 8 years of experience in providing comprehensive support and deployment services, resolving issues for hardware and software products.
• Proficient knowledge and experience in Linux system administration, including process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging, and advanced networking (tuning and monitoring).
• Familiarity with cluster management and provisioning technologies for bare-metal servers (bonus for experience with BCM (Base Command Manager)).
• A minimum of a four-year degree from an accredited university or college in Computer Science, Electrical or Computer Engineering, or equivalent experience.
• Proficiency in scripting languages such as Bash, Python, Ansible, etc.
• Exceptional interpersonal skills with a proven ability to resolve customer issues as they arise.
• Strong organizational capabilities and adeptness at prioritizing and multi-tasking with minimal supervision.
• Experience with job schedulers such as SLURM, LSF, UGE, etc.
• Willingness to travel to customer locations within the United States up to 20% of the time.
• Familiarity with benchmarking tools like HPL, NCCL tests, MLPerf, along with experience in Kubernetes.
• Knowledge of InfiniBand technology.
• Experience with GPU (Graphics Processing Unit) focused hardware and software.
• Experience with MPI (Message Passing Interface).
• Knowledge of storage technologies such as Lustre or GPFS.
• Familiarity with OEM GPU platforms.
• Equity.
• Comprehensive benefits.
Unconventional AI
Scale Army Careers
TubeScience
Get handpicked remote jobs straight to your inbox weekly.