
AI Compute Engineer
Posted 18 hours ago

Posted 18 hours ago
This is a fully remote position, open to applicants in California.
• Implement, oversee, and validate AI Compute/HPC infrastructure within Linux-based environments for both new and existing clients.
• Serve as the subject matter expert, guiding customers from initial planning discussions through to successful implementation.
• Collaborate with customers, partners, and internal teams to analyze, define, and execute extensive AI Compute projects.
• Engage across networking, system design, automation, and validation processes.
• Create handover documentation and conduct knowledge transfers as clients initiate the deployment of advanced systems.
• Provide internal teams with insights by reporting bugs, documenting workarounds, and recommending enhancements.
• Over 4 years of experience delivering comprehensive support and deployment services, along with problem-solving for hardware and software products.
• Proficient knowledge and experience in Linux system administration, including process management, package management, task scheduling, kernel management, boot procedures/troubleshooting, performance reporting/optimization/logging, and advanced networking.
• Familiarity with cluster management and provisioning technologies for bare-metal servers; additional credit for experience with BCM (Base Command Manager).
• A minimum of a four-year degree from an accredited university or college in Computer Science, Electrical or Computer Engineering, or equivalent experience.
• Proficient scripting skills in Bash, Python, Ansible, or similar languages.
• Exceptional interpersonal skills with the ability to resolve customer issues effectively.
• Strong organizational skills, capable of prioritizing and managing multiple tasks with minimal supervision.
• Experience with scheduling systems such as SLURM, LSF, or UGE.
• Willingness to travel to customer locations within the United States up to 20% of the time.
• Familiarity with benchmarking tools such as HPL, NCCL tests, and MLPerf.
• Experience with Kubernetes.
• Knowledge of InfiniBand.
• Experience with hardware/software focused on GPU technology.
• Familiarity with MPI.
• Experience with storage technologies like Lustre or GPFS.
• Awareness of OEM GPU platforms.
• Equity
• Benefits
• Equal opportunity employer
Jazz Pharmaceuticals
Hydra Host
Grupo Consciência
RWS Group
Get handpicked remote jobs straight to your inbox weekly.