
DevOps, GitLab-based Platform, CICD 30*3 Pipelines
Posted Aug 25

Posted Aug 25
This is a fully remote position, open to applicants in California, +1 more state.
• Design, develop, and sustain comprehensive CI/CD pipelines for software applications and machine learning models.
• Automate processes for building, testing, deploying, and rolling back software.
• Create, optimize, and expand cloud-native infrastructure utilizing Kubernetes and Docker.
• Oversee and allocate specialized computing resources, including GPU clusters, to support high-performance AI workloads and model inferencing.
• Take responsibility for high-availability design in production settings.
• Develop and implement disaster recovery plans, self-healing strategies, capacity forecasting, and performance optimization.
• Advocate for Infrastructure as Code methodologies using Terraform, Ansible, and Helm.
• Design and enhance monitoring, logging, and alerting systems with tools like Prometheus, Grafana, and the ELK/EFK stack.
• Construct the Internal Developer Platform and streamlined processes that allow product, model, and data-science teams to deploy without ticket submissions.
• Work alongside R&D, Data Science, Security, and Business teams to improve workflows and eliminate inefficiencies.
• Set and uphold standards for system stability and security.
• Manage release workflows, implement Zero Trust access controls, supervise secrets management, and ensure compliance with SOC2/ISO27001.
• Lead troubleshooting efforts, root-cause analysis, and proactive remediation during complex issues and major incidents.
• Transform incident insights into automation to avert future occurrences.
• Bachelor's degree or higher in Computer Science, Engineering, or a related technical discipline.
• Over 5 years of practical experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles.
• Advanced knowledge of Linux operating systems and fundamental networking concepts, including TCP/IP, DNS, HTTP, Load Balancing, and VPCs.
• Profound expertise in Docker and Kubernetes orchestration, including cluster management and best practices for production environments.
• Proficient in designing and managing infrastructure on major public or hybrid cloud platforms, such as AWS, GCP, Azure, or Alibaba Cloud.
• Experience with multi-cloud and hybrid-cloud approaches.
• Strong programming and scripting skills in at least one prominent language, such as Go, Python, or Shell.
• Methodical and practical knowledge of CI/CD practices, Infrastructure as Code (IaC), observability frameworks, and SRE principles.
• Outstanding problem-solving skills and keen technical judgment.
• Excellent communication abilities across teams.
• Preferred: Familiarity with MLOps practices, model serving/inferencing frameworks like vLLM, TGI, or Triton Inference Server.
• Preferred: Experience managing GPU clusters for AI/ML tasks.
• Preferred: Background in large-scale distributed systems or high-concurrency environments.
• Preferred: Practical experience in designing and developing Internal Developer Platforms (IDP).
• Preferred: Knowledge of Zero Trust architecture, automated security testing (DevSecOps), SOC2, or ISO27001.
• Preferred: Previous experience as a Technical Lead, mentoring junior engineers, or leading DevOps teams.
• Preferred: Experience integrating an LLM-driven code/config helper into a pipeline or strong perspectives on the subject.
• Must adhere to applicable work authorization and equal employment regulations in the relevant country, state, and local jurisdictions.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and the option for remote work.
• Opportunities for professional development and continuous learning.
• Generous paid time off and holiday schedule.
• Supportive and inclusive company culture.
• Access to cutting-edge tools and technologies.
Bet On Talent
Virtasant
Ookla
opinov8
Get handpicked remote jobs straight to your inbox weekly.