
Senior Site Reliability Engineer, Data & Analytics
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Participate in an on-call rotation and drive incidents to resolution
• Lead blameless postmortems and identify systemic improvements for reliability
• Collaborate with data, ML, and platform teams on batch, streaming, training, and inference workloads
• Support ML training pipelines and inference services, including those requiring GPU workloads
• Help define the operational framework for data and ML services on Kubernetes
• Design and develop automation and operational tools, including workflows, diagnostic tools, and runbooks
• Build and enhance centralized platform services, shared tools, data integrations, and access controls
• Diagnose and resolve issues related to reliability, performance, and costs across distributed systems
• Advocate for automation, documentation, and practices that minimize toil
• Maintain infrastructure using Terraform and adhere to infrastructure-as-code principles
• Enhance CI/CD and GitOps workflows utilizing Jenkins, GitHub Actions, and ArgoCD
• Operate and improve containerized services within Kubernetes
• Define and measure reliability metrics using SLIs, SLOs, and error budgets
• Conduct load testing, capacity modeling, and production validation
• Develop internal tools and streamlined processes that enable teams to operate safely and efficiently
• Experience in operating reliable, distributed systems in SRE, platform, or similar roles
• Familiarity with data, analytics, ML, or large-scale distributed workloads
• Strong understanding of Linux, containers, Kubernetes, and cloud infrastructure
• Proficient in building automation or internal tools using Python, Go, shell, or similar languages
• Experience with infrastructure-as-code practices, such as Terraform
• Knowledge of CI/CD or GitOps systems, including Jenkins, GitHub Actions, or ArgoCD
• Familiarity with observability tools, including metrics, logs, traces, alerting, and incident response
• Solid grasp of SRE principles, including SLIs, SLOs, error budgets, and postmortems
• Proven track record of employing modern development and automation practices to enhance reliability and efficiency
• Experience in creating internal tooling, automation, or systems that boost developer productivity
• Strong communication skills for collaborating with technical and cross-functional teams
• Bonus: Experience with data and ML systems, such as training pipelines, model serving, and GPU workloads
• Bonus: Experience with distributed systems and messaging technologies, such as Kafka or Pub/Sub
• Bonus: Experience in Kubernetes-based environments
• Bonus: Familiarity with Prometheus and Grafana
• Bonus: Experience operating systems within GCP or AWS cloud environments
• Medical, dental, and vision coverage
• Health savings account or health reimbursement account
• Healthcare spending accounts
• Dependent care spending accounts
• Life and AD&D insurance
• Disability insurance
• 401(k) with company match
• Tuition reimbursement
• Charitable donation matching
• Paid holidays and vacation
• Paid sick time
• Floating holidays
• Compassion and bereavement leaves
• Parental leave
• Mental health and wellbeing programs
• Fitness programs
• Free and discounted games
• Supplemental life and disability benefits
• Legal service
• ID protection
• Rental insurance
• Relocation assistance may be available if the company requires geographic relocation
• Incentive compensation may be available
Keyfactor
Trumid
GSB Solutions
Smartcat
Get handpicked remote jobs straight to your inbox weekly.