Remotery

Senior DevOps Engineer – AI Cloud

Posted 1 day ago

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Design, implement, and maintain comprehensive CI/CD pipelines for software applications and machine learning models.

• Automate processes for building, testing, deploying, and rolling back applications.

• Develop, optimize, and scale cloud-native infrastructure utilizing Kubernetes and Docker.

• Manage and provision specialized computing resources, including GPU clusters, to support AI workloads and model inferencing.

• Take ownership of high-availability designs in production environments.

• Implement disaster recovery strategies, self-healing mechanisms, capacity planning, and performance tuning.

• Advocate for Infrastructure as Code practices using tools like Terraform, Ansible, and Helm.

• Architect and enhance monitoring, logging, and alerting systems.

• Collaborate with Research & Development, Data Science, Security, and Business teams on optimizing workflows and advancing Platform Engineering initiatives.

• Establish and uphold standards for system stability and security, release workflows, Zero Trust access controls, secrets management, and compliance.

• Lead troubleshooting efforts during complex anomalies and major incidents, perform root cause analysis, and develop preventative remediation plans.


⛳️ Requirements

• Bachelor's degree or higher in Computer Science, Engineering, or a related technical field.

• Over 5 years of practical experience in DevOps, Site Reliability Engineering, or Cloud Infrastructure positions.

• Expert knowledge of Linux operating systems and fundamental networking concepts, including TCP/IP, DNS, HTTP, load balancing, and VPCs.

• Advanced expertise in Docker and Kubernetes orchestration, cluster management, and production best practices.

• Proficient in designing and managing infrastructure on leading public or hybrid cloud platforms, including AWS, GCP, Azure, or Alibaba Cloud.

• Experience with multi-cloud and hybrid-cloud strategies.

• Strong coding and scripting skills in at least one major programming language, such as Go, Python, or Shell.

• Solid understanding of CI/CD, Infrastructure as Code, observability, and Site Reliability Engineering principles.

• Outstanding problem-solving skills, technical judgment, and the ability to communicate effectively across teams.

• Preferred: familiarity with MLOps, model serving/inferencing frameworks, GPU clusters, large-scale distributed systems, Internal Developer Platforms, Zero Trust, DevSecOps, SOC2, ISO27001, and technical leadership.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible work hours and opportunities for remote work.

• Professional development and training programs.

• Collaborative and innovative work environment.

People also viewed

CVS Health6 hours ago

Salesforce DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$83.4k – $166.9k/year
ApplyView job
Devoteam7 hours ago

Data, AWS DevSecOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aspirion7 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Goodgame Studios7 hours ago

Senior Agentic Engineer – Java Backend, DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Instacart8 hours ago

Site Reliability Engineer II

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$133k – $169k/year
ApplyView job
Logicalis Spain8 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€40k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers