Remotery

Site Reliability Engineer

atRunPodRemoteUS flagCaliforniaFull-timeDevOps & Site Reliability Engineer (SRE)Mid-levelSenior$150k – $200k/year

Posted Jul 28

This is a fully remote position, open to applicants in California.

📋 Description

• Establish and execute SLIs/SLOs for essential services.

• Direct incident response and oversee cross-team mitigation initiatives.

• Carry out blameless postmortems and ensure all corrective measures are implemented.

• Conduct production readiness assessments for new services and features.

• Recognize systemic risks and promote preventive enhancements.

• Design and enhance monitoring systems, alerting mechanisms, and dashboards (Prometheus, Grafana, etc.).

• Develop internal tools for tracking and reporting reliability.

• Automate recurring operational tasks.

• Enhance CI/CD reliability along with release processes.

• Collaborate with engineering teams to bolster system resilience.

• Offer expertise on fault tolerance, scalability, and failure management.

• Engage in architectural discussions with a focus on reliability.


⛳️ Requirements

• Over 5 years of experience in SRE, Reliability Engineering, or Production Engineering.

• Proficient in Linux systems and networking.

• Experience overseeing containerized production environments.

• Strong knowledge of distributed systems and their failure modes.

• Proven ability to define and manage SLIs/SLOs.

• Demonstrated experience in incident response and leading postmortems.

• Strong skills in scripting or programming.

• Familiarity with monitoring and alerting systems.

• Outstanding written communication abilities.

• Successful completion of a background check.

• Preferred: Experience with GPU infrastructure or AI/ML platforms.

• Experience enhancing reliability in high-growth or large-scale settings.

• Knowledge of GPU observability tools.

• Experience with Infrastructure as Code.

• Background in startup environments.

• Experience creating internal reliability platforms or frameworks.


🏝️ Benefits

• Meaningful equity in a rapidly growing company—everyone on the team receives stock options, allowing you to share in the company's success as your contributions drive growth.

• Comprehensive medical, dental, and vision plans.

• Flexible PTO—take the time you need to rejuvenate.

• Most positions are remote-first, with an inclusive and collaborative team using Slack as the primary communication tool.

• Become part of a passionate team at the forefront of AI infrastructure—where culture, learning, and ownership are central to our scaling efforts.

People also viewed

TEKsystems16 hours ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems16 hours ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data19 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job
Level Data19 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data19 hours ago

DevOps Engineer II

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$95k – $110k/year
ApplyView job
Level Data19 hours ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers