Remotery

Staff Site Reliability Engineer

Posted Jun 23

This is a fully remote position, open to applicants in United States.

📋 Description

• Design Reliability Frameworks: Create frameworks and self-service tools that empower teams to take ownership of their service reliability within a “You Build It, You Run It” philosophy.

• Lead AI-Enhanced Reliability: Spearhead our AIOps initiatives by automating diagnostics, remediation processes, and proactive failure prevention.

• Promote a Reliability Culture: Integrate Site Reliability Engineering (SRE) practices throughout the engineering process via design evaluations, production readiness checks, and operational benchmarks.

• Incident Management Leadership: Serve as Incident Commander during critical incidents, exemplifying operational excellence and ensuring that blameless postmortems yield sustainable improvements.

• Enhance Observability: Provide comprehensive monitoring, tracing, and profiling (using tools like Prometheus, Grafana, OTEL, Continuous Profiling) to proactively enhance performance.

• Mentor and Scale: Elevate the capabilities of engineers within SRE and product teams through mentorship, technical direction, and knowledge dissemination.


⛳️ Requirements

• 8+ years of experience in Site Reliability Engineering, DevOps, or a related field, including a minimum of 3+ years in a Senior+ SRE role.

• Strong expertise in managing production SaaS systems at scale.

• Proficient in at least one programming or scripting language (Python, Go, or similar).

• Practical experience with cloud services (AWS, GCP, or Azure) and Kubernetes.

• Comprehensive understanding of networking principles (TCP/IP, DNS, HTTP/S, load balancing).

• Familiarity with monitoring and alerting tools (Prometheus, Grafana, Datadog, ELK).

• Knowledge of advanced observability tools (OTEL, continuous profiling).

• Demonstrated incident management proficiency, particularly in leading high-severity incidents and conducting postmortems.

• Strong troubleshooting abilities across the entire technology stack.

• Exceptional communication and collaboration skills.


🏝️ Benefits

• You may also be offered equity.

• A generous benefits program.

People also viewed

The CodestJul 26

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
IRIUMJul 26

Ingeniero/a Cloud DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€33k – €40k/year
ApplyView job
SólidesJul 26

Senior DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ResilincJul 25

Junior/Senior Site Reliability Engineer – Night Shift

IN flagIndia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity GroupJul 25

Senior SRE / DevOps Engineer

Anywhere in the WorldFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
HOESSLER & HOESSLERJul 25

DevOps Software Engineer – Career Ambitions

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€65k – €75k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers