
Site Reliability Engineer – ML Infrastructure
Posted Jul 14

Posted Jul 14
This is a fully remote position, open to applicants in Japan.
• Adhering to SRE principles to ensure a 24/7 production environment utilizing Kubernetes.
• Implementing DevOps methodologies to enhance the quality of life for the IT team.
• Conducting proactive system monitoring and configuration management.
• Executing incident response and postmortem analysis processes.
• Overseeing and advancing AWS infrastructure components, including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
• Developing and maintaining CI/CD pipelines alongside infrastructure as code practices using Terraform, Helm, and ArgoCD.
• Guaranteeing system reliability, performance, and scalability across our production environment.
• Implementing SRE principles to machine learning infrastructure, ensuring that model serving, training pipelines, and data systems are reliable, observable, and effectively managed.
• Enhancing ML model deployment pipelines and MLOps practices.
• Monitoring the performance of ML models in production while establishing alerting and observability for ML systems.
• Collaborating with data scientists and product teams to operationalize ML models at scale.
• Contributing to the infrastructure supporting ML workloads on Kubernetes and AWS.
• A minimum of 2 years of experience with Amazon Web Services (AWS), particularly focusing on EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.
• Substantial hands-on experience with AWS EKS.
• Direct software engineering experience following DevOps/SRE practices, with at least 1 year in a technical lead role.
• Proficiency in at least one of the following programming languages: Python, Ruby, Elixir, Go, Javascript, or Rust.
• Understanding of container and hypervisor fundamentals.
• Experience with configuration management (YAML/Bash); familiarity with Helm and Terraform is preferred.
• Proven experience running production systems at a large scale, with an understanding of potential issues and their solutions.
• Knowledge of machine learning workflows and MLOps practices.
• Python experience with ML-related tools (model deployment, inference serving, or ML pipeline tooling).
• Fully remote working environment.
Ontrac Solutions
CyberSheath
Ontrac Solutions
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.