
Senior Site Reliability Engineer, SRE
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
• Drive initiatives to enhance system reliability, scalability, and performance across essential services.
• Establish and execute SLIs, SLOs, and error budgets to inform engineering priorities.
• Design and implement observability frameworks encompassing metrics, logging, tracing, and alerting.
• Lead intricate incident response efforts and serve as the incident commander when necessary.
• Conduct postmortems that focus on systemic issues and ensure that corrective measures are executed.
• Identify and eliminate repetitive tasks through automation, tooling, and optimized workflows.
• Collaborate with product and platform teams regarding architectural decisions, production readiness, and failure recovery strategies.
• Create reusable systems and paved pathways for dependable service operations.
• Mentor engineers to elevate organizational operational maturity.
• Achieve well-defined SLOs, actionable alerts, efficient incident management, enhanced adoption of reliability practices, and decreased toil.
• 6-10+ years of experience in Site Reliability Engineering (SRE), infrastructure, or backend systems engineering.
• Proven track record of owning reliability outcomes for intricate, distributed systems.
• Extensive experience with cloud infrastructure (AWS, GCP, or Azure) and production-scale systems.
• In-depth understanding of observability, incident management, and system performance.
• Proficient in at least one programming language, such as Go, Python, or Java, with a focus on automation and tooling.
• Capability to influence how other teams operate without holding managerial authority.
• Ability to make decisive choices during incidents by adhering to a defined process while remaining composed.
• Legally authorized to work in the United States without current or future employer-sponsored visa sponsorship.
• Adherence to UJET's legal, regulatory, security, data protection, and policy requirements.
• Notable qualifications: SRE practices, Kubernetes/container orchestration, Infrastructure as Code (IaC) such as Terraform, experience with high-growth or scaling systems, and expertise in performance engineering or capacity planning.
• Medical insurance
• Dental insurance
• Vision insurance
• 401(k) plan
• Wellness benefits
• Meaningful work that shapes the future of customer experience
• A collaborative and inclusive team culture
• Equal employment opportunities
• SDPC training
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.