
Senior DevOps Engineer – AI Cloud
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California, +1 more state.
• Design, implement, and maintain comprehensive CI/CD pipelines for software applications and machine learning models.
• Automate processes for building, testing, deploying, and rolling back applications.
• Develop, optimize, and scale cloud-native infrastructure utilizing Kubernetes and Docker.
• Manage and provision specialized computing resources, including GPU clusters, to support AI workloads and model inferencing.
• Take ownership of high-availability designs in production environments.
• Implement disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning.
• Advocate for Infrastructure as Code practices using tools like Terraform, Ansible, and Helm.
• Architect and enhance monitoring, logging, and alerting systems.
• Collaborate with Research & Development, Data Science, Security, and Business teams on optimizing workflows and advancing Platform Engineering initiatives.
• Establish and uphold standards for system stability and security, release workflows, Zero Trust access controls, secrets management, and compliance.
• Lead troubleshooting efforts during complex anomalies and major incidents, perform root cause analysis, and develop preventative remediation plans.
• Bachelor's degree or higher in Computer Science, Engineering, or a related technical field.
• Over 5 years of practical experience in DevOps, Site Reliability Engineering, or Cloud Infrastructure positions.
• Expert knowledge of Linux operating systems and fundamental networking concepts, including TCP/IP, DNS, HTTP, load balancing, and VPCs.
• Advanced expertise in Docker and Kubernetes orchestration, cluster management, and production best practices.
• Proficient in designing and managing infrastructure on leading public or hybrid cloud platforms, including AWS, GCP, Azure, or Alibaba Cloud.
• Experience with multi-cloud and hybrid-cloud strategies.
• Strong coding and scripting skills in at least one major programming language, such as Go, Python, or Shell.
• Solid understanding of CI/CD, Infrastructure as Code, observability, and Site Reliability Engineering principles.
• Outstanding problem-solving skills, technical judgment, and the ability to communicate effectively across teams.
• Preferred: familiarity with MLOps, model serving/inferencing frameworks, GPU clusters, large-scale distributed systems, Internal Developer Platforms, Zero Trust, DevSecOps, SOC2, ISO27001, and technical leadership.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work hours and opportunities for remote work.
• Professional development and training programs.
• Collaborative and innovative work environment.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.