Senior Site Reliability Engineer, SRE

Posted 3 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Drive initiatives to enhance system reliability, scalability, and performance across essential services.

• Establish and execute SLIs, SLOs, and error budgets to inform engineering priorities.

• Design and implement observability frameworks encompassing metrics, logging, tracing, and alerting.

• Lead intricate incident response efforts and serve as the incident commander when necessary.

• Conduct postmortems that focus on systemic issues and ensure that corrective measures are executed.

• Identify and eliminate repetitive tasks through automation, tooling, and optimized workflows.

• Collaborate with product and platform teams regarding architectural decisions, production readiness, and failure recovery strategies.

• Create reusable systems and paved pathways for dependable service operations.

• Mentor engineers to elevate organizational operational maturity.

• Achieve well-defined SLOs, actionable alerts, efficient incident management, enhanced adoption of reliability practices, and decreased toil.


⛳️ Requirements

• 6-10+ years of experience in Site Reliability Engineering (SRE), infrastructure, or backend systems engineering.

• Proven track record of owning reliability outcomes for intricate, distributed systems.

• Extensive experience with cloud infrastructure (AWS, GCP, or Azure) and production-scale systems.

• In-depth understanding of observability, incident management, and system performance.

• Proficient in at least one programming language, such as Go, Python, or Java, with a focus on automation and tooling.

• Capability to influence how other teams operate without holding managerial authority.

• Ability to make decisive choices during incidents by adhering to a defined process while remaining composed.

• Legally authorized to work in the United States without current or future employer-sponsored visa sponsorship.

• Adherence to UJET's legal, regulatory, security, data protection, and policy requirements.

• Notable qualifications: SRE practices, Kubernetes/container orchestration, Infrastructure as Code (IaC) such as Terraform, experience with high-growth or scaling systems, and expertise in performance engineering or capacity planning.


🏝️ Benefits

• Medical insurance

• Dental insurance

• Vision insurance

• 401(k) plan

• Wellness benefits

• Meaningful work that shapes the future of customer experience

• A collaborative and inclusive team culture

• Equal employment opportunities

• SDPC training

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers