Senior Site Reliability Engineer, Data & Analytics

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Participate in an on-call rotation and drive incidents to resolution

• Lead blameless postmortems and identify systemic improvements for reliability

• Collaborate with data, ML, and platform teams on batch, streaming, training, and inference workloads

• Support ML training pipelines and inference services, including those requiring GPU workloads

• Help define the operational framework for data and ML services on Kubernetes

• Design and develop automation and operational tools, including workflows, diagnostic tools, and runbooks

• Build and enhance centralized platform services, shared tools, data integrations, and access controls

• Diagnose and resolve issues related to reliability, performance, and costs across distributed systems

• Advocate for automation, documentation, and practices that minimize toil

• Maintain infrastructure using Terraform and adhere to infrastructure-as-code principles

• Enhance CI/CD and GitOps workflows utilizing Jenkins, GitHub Actions, and ArgoCD

• Operate and improve containerized services within Kubernetes

• Define and measure reliability metrics using SLIs, SLOs, and error budgets

• Conduct load testing, capacity modeling, and production validation

• Develop internal tools and streamlined processes that enable teams to operate safely and efficiently


⛳️ Requirements

• Experience in operating reliable, distributed systems in SRE, platform, or similar roles

• Familiarity with data, analytics, ML, or large-scale distributed workloads

• Strong understanding of Linux, containers, Kubernetes, and cloud infrastructure

• Proficient in building automation or internal tools using Python, Go, shell, or similar languages

• Experience with infrastructure-as-code practices, such as Terraform

• Knowledge of CI/CD or GitOps systems, including Jenkins, GitHub Actions, or ArgoCD

• Familiarity with observability tools, including metrics, logs, traces, alerting, and incident response

• Solid grasp of SRE principles, including SLIs, SLOs, error budgets, and postmortems

• Proven track record of employing modern development and automation practices to enhance reliability and efficiency

• Experience in creating internal tooling, automation, or systems that boost developer productivity

• Strong communication skills for collaborating with technical and cross-functional teams

• Bonus: Experience with data and ML systems, such as training pipelines, model serving, and GPU workloads

• Bonus: Experience with distributed systems and messaging technologies, such as Kafka or Pub/Sub

• Bonus: Experience in Kubernetes-based environments

• Bonus: Familiarity with Prometheus and Grafana

• Bonus: Experience operating systems within GCP or AWS cloud environments


🏝️ Benefits

• Medical, dental, and vision coverage

• Health savings account or health reimbursement account

• Healthcare spending accounts

• Dependent care spending accounts

• Life and AD&D insurance

• Disability insurance

• 401(k) with company match

• Tuition reimbursement

• Charitable donation matching

• Paid holidays and vacation

• Paid sick time

• Floating holidays

• Compassion and bereavement leaves

• Parental leave

• Mental health and wellbeing programs

• Fitness programs

• Free and discounted games

• Supplemental life and disability benefits

• Legal service

• ID protection

• Rental insurance

• Relocation assistance may be available if the company requires geographic relocation

• Incentive compensation may be available

People also viewed

Keyfactor1 day ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Trumid1 day ago

Senior Database Reliability Engineer – DBRE

US flagNew York OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$225k – $265k/year
ApplyView job
GSB Solutions1 day ago

Senior Release Engineer

MX flagMexico OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$80k/year
ApplyView job
Smartcat1 day ago

Head of Infrastructure and DevOps

PT flagPortugal OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Smartcat1 day ago

Head of Infrastructure and DevOps

GE flagGeorgia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Physitrack PLC1 day ago

Mid DevOps Engineer – Security, Reliability

PL flagPoland OnlyFreelanceDevOps & Site Reliability Engineer (SRE)€6,000 – €8,000/month
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers