
Software Engineer – AI Research Clusters
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in California, +5 more states.
• Propose and execute engineering solutions for functional, reliable, secure, and performance-optimized GPU clusters
• Minimize operational disruption and overhead for internal AI researchers
• Facilitate self-service continuous improvement in reliability, operational excellence, and performance
• Identify pain points in validating, monitoring, and managing GPU clusters at scale
• Design, develop, and sustain engineering solutions to tackle cluster operations challenges
• Investigate traditional AIOps and emerging Agentic AI to alleviate operational toil
• Engage in on-call support for systems and platforms created and maintained by the team
• Collaborate with colleagues throughout the AI Platform organization
• BS/MS in Computer Science, Engineering, or equivalent experience
• 2+ years in software/platform engineering, with at least 1 year focused on ML infrastructure or distributed systems
• Experience in the software development lifecycle on Linux-based platforms
• Proficient coding skills in Python, C++, or Rust
• Familiarity with Docker, Kubernetes, GitLab CI, and automated deployments
• Experience with AIOps or Agentic AI and successfully applying it in a production environment
• Expertise in full-stack development, including relational data modeling, database optimization, REST API semantics, JavaScript, CSS, and providing API as a service
• Experience operating Slurm or custom scheduling frameworks in production ML environments
• Knowledge of GPU computing, Linux systems internals, and performance tuning at scale
• Equity
• Benefits
Sardine
accesa.eu
Aptura
Thermo Fisher Scientific
Get handpicked remote jobs straight to your inbox weekly.