
Cloud Service Security DevOps – Maintenance
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in California, +1 more state.
• Design, develop, and sustain comprehensive CI/CD pipelines for software applications and machine learning models.
• Automate processes related to building, testing, deployment, and rollback.
• Create, enhance, and scale cloud-native infrastructure utilizing Kubernetes and Docker.
• Oversee and provision GPU clusters and specialized computing resources tailored for AI workloads and model inferencing.
• Take ownership of high-availability architecture within production environments.
• Implement disaster recovery solutions, self-healing mechanisms, capacity planning, and performance optimization.
• Develop infrastructure provisioning that is automated, reproducible, and auditable across various cloud environments.
• Architect and improve systems for monitoring, logging, and alerting.
• Construct the Internal Developer Platform and establish optimal deployment paths for product, model, and data-science teams.
• Collaborate effectively with R&D, Data Science, Security, and Business teams.
• Set and uphold standards for system stability and security.
• Manage release workflows, implement Zero Trust access controls, handle secrets management, and ensure compliance.
• Lead troubleshooting efforts and root-cause analysis during complex anomalies and significant incidents.
• Implement preventive remediation strategies and automate responses to recurring incidents.
• A Bachelor's degree or higher in Computer Science, Engineering, or a related technical discipline.
• Over 5 years of practical experience in DevOps, SRE, or Cloud Infrastructure positions.
• Expert knowledge of Linux operating systems and essential networking principles, including TCP/IP, DNS, HTTP, load balancing, and VPCs.
• Extensive expertise in Docker and Kubernetes orchestration, cluster management, and production best practices.
• Proficient in designing and overseeing infrastructure on major public or hybrid cloud platforms such as AWS, GCP, Azure, or Alibaba Cloud.
• Experience with multi-cloud and hybrid-cloud strategies.
• Strong coding and scripting skills in Go, Python, Shell, or another prominent programming language.
• Understanding of CI/CD, Infrastructure as Code, observability, and SRE principles.
• Outstanding problem-solving skills and technical judgment.
• Excellent communication skills across teams.
• Preferred: familiarity with MLOps and model serving/inferencing frameworks like vLLM, TGI, or Triton Inference Server.
• Preferred: experience managing GPU clusters for AI/ML workloads.
• Preferred: experience with large-scale distributed systems or high-concurrency environments.
• Preferred: experience in designing and building Internal Developer Platforms.
• Preferred: familiarity with Zero Trust architecture, DevSecOps, SOC2, and ISO27001.
• Preferred: experience as a Technical Lead, mentoring junior engineers, or managing DevOps teams.
• Competitive salary and performance-based incentives.
• Opportunities for professional growth and development.
• Flexible work environment with remote work options.
• Comprehensive health, dental, and vision insurance.
• Generous vacation and paid time off policies.
Guidehouse
GFT Technologies
SysMap Solutions
CENTERLIGHT
Get handpicked remote jobs straight to your inbox weekly.