Remotery

Senior Site Reliability Engineer

Posted Jul 28

This is a fully remote position, open to applicants in California.

📋 Description

• Act as the primary responder in cases of cluster outages or issues, promptly triaging and resolving urgent matters as they occur.

• Maintain a high level of cluster uptime (measured in multiple nines) and establish plus monitor SLAs to assess reliability.

• Identify systemic and recurring issues, engineering precise solutions in partnership with engineering teams.

• Create strong metrics and observability for cluster health, utilizing these metrics to guide your efforts. Develop custom observability tools when existing solutions are inadequate.

• Collaborate with software and research teams to establish policies for equitable cluster usage, and assist in developing mechanisms to enforce these policies.

• Aid in predicting cluster growth and selecting suitable scale-up strategies, while optimizing operations for cost and usability.


⛳️ Requirements

• Over 5 years of experience in SRE or DevOps positions, ideally in a senior engineer or technical lead role.

• Proficient in HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod).

• Ability to create scripts and utilities of moderate complexity using a common scripting language (Python, Ruby, etc.).

• Familiar with infrastructure-as-code and configuration management tools (Terraform, Ansible).

• Experience with cloud infrastructure platforms (AWS or GCP).

• Knowledge of designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry).

• Experience with distributed storage technologies (Lustre, Ceph, S3).

• Possess a "system engineer" mindset rather than a "system administrator" perspective, employing systematic thinking and automation.

• Bachelor's degree in computer science.


🏝️ Benefits

• Medical, dental, and vision coverage.

• Life and AD&D insurance.

• 20 days of paid time off.

• 9 sick days.

• 401(k) plan with company matching.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers