Remotery

Site Reliability Engineer – ML Infrastructure

Posted Jul 14

This is a fully remote position, open to applicants in Japan.

📋 Description

• Adhering to SRE principles to ensure a 24/7 production environment utilizing Kubernetes.

• Implementing DevOps methodologies to enhance the quality of life for the IT team.

• Conducting proactive system monitoring and configuration management.

• Executing incident response and postmortem analysis processes.

• Overseeing and advancing AWS infrastructure components, including EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.

• Developing and maintaining CI/CD pipelines alongside infrastructure as code practices using Terraform, Helm, and ArgoCD.

• Guaranteeing system reliability, performance, and scalability across our production environment.

• Implementing SRE principles to machine learning infrastructure, ensuring that model serving, training pipelines, and data systems are reliable, observable, and effectively managed.

• Enhancing ML model deployment pipelines and MLOps practices.

• Monitoring the performance of ML models in production while establishing alerting and observability for ML systems.

• Collaborating with data scientists and product teams to operationalize ML models at scale.

• Contributing to the infrastructure supporting ML workloads on Kubernetes and AWS.


⛳️ Requirements

• A minimum of 2 years of experience with Amazon Web Services (AWS), particularly focusing on EKS, EC2, RDS, Fargate, CloudFront, Lambda, and S3.

• Substantial hands-on experience with AWS EKS.

• Direct software engineering experience following DevOps/SRE practices, with at least 1 year in a technical lead role.

• Proficiency in at least one of the following programming languages: Python, Ruby, Elixir, Go, Javascript, or Rust.

• Understanding of container and hypervisor fundamentals.

• Experience with configuration management (YAML/Bash); familiarity with Helm and Terraform is preferred.

• Proven experience running production systems at a large scale, with an understanding of potential issues and their solutions.

• Knowledge of machine learning workflows and MLOps practices.

• Python experience with ML-related tools (model deployment, inference serving, or ML pipeline tooling).


🏝️ Benefits

• Fully remote working environment.

People also viewed

Ontrac Solutions2 days ago

Site Reliability Engineer

PK flagPakistan OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
CyberSheath2 days ago

Cloud Operations Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$110k – $127k/year
ApplyView job
Ontrac Solutions2 days ago

Site Reliability Engineer

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
NVIDIA2 days ago

Service Reliability Engineer

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$168k – $333.5k/year
ApplyView job
Nagarro2 days ago

Senior Site Reliability Engineer, AWS Cloud

RO flagRomania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini2 days ago

Senior DevOps Engineer

UA flagUkraine OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers