Senior Site Reliability Engineer – Infra Ops

Posted 3 days ago

This is a fully remote position, open to applicants in California.

📋 Description

• Design, construct, and manage Kubernetes platforms that ensure secure, highly available, and scalable critical production services across hybrid and public cloud settings.

• Develop infrastructure as code using Terraform with reusable modules, secure delivery workflows, and controlled infrastructure modifications.

• Create backend services, internal tools, and operational automation utilizing Go, Python, or JavaScript/TypeScript.

• Collaborate with engineering and product teams to convert workload requirements into effective solutions focusing on reliability, performance, capacity, security, and cost.

• Enhance the production lifecycle through CI/CD, deployment automation, progressive delivery, and clearly defined operational ownership.

• Establish and refine observability practices spanning metrics, logs, traces, alerting, and dashboards.

• Engage in on-call duties, lead incident response, conduct root-cause analysis, and facilitate blameless postmortems and corrective actions.

• Set reliability benchmarks using SLIs, SLOs, error budgets, capacity planning, disaster-recovery tests, and resilience improvements.

• Integrate security and compliance measures into platform operations.

• Utilize AI-assisted and data-driven operational strategies to enhance signal detection, minimize alert noise, expedite root-cause analysis, and identify opportunities for automation.

• Actively participate in code reviews, documentation, knowledge sharing, and mentoring.

• Provide mentorship and support for team development.


⛳️ Requirements

• A minimum of 5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a similar software engineering role focused on production systems.

• Extensive hands-on experience with Kubernetes, including the design, operation, security, and troubleshooting of production clusters and containerized workloads at scale.

• Proficient in Terraform, including creating reusable modules, managing state and environments, and implementing infrastructure changes through automated, reviewable workflows.

• Experience in production software development using Go, Python, or JavaScript/TypeScript.

• Proven track record of enhancing the reliability, performance, scalability, or cost-effectiveness of distributed systems in live environments.

• Familiarity with cloud infrastructure and essential networking concepts, such as IAM, DNS, load balancing, routing, service networking, and secure connectivity.

• Strong skills in observability and troubleshooting utilizing metrics, logs, traces, alerting, and incident data.

• Experience in defining and operating based on SLIs, SLOs, error budgets, incident management processes, postmortems, and disaster recovery practices.

• Knowledge of CI/CD, GitOps or deployment automation, as well as canary or blue-green deployments.

• A security-focused mindset regarding infrastructure and experience collaborating with security and engineering teams in regulated or high-availability environments.

• Excellent written and verbal communication skills, along with a strong sense of ownership and judgment to balance speed, risk, and operational excellence.

• Experience with AI-assisted tools in engineering or operational workflows is an advantage.


🏝️ Benefits

• Flexible work environment.

• Remote-first work arrangement.

People also viewed

Koniag Government Services2 days ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)2 days ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus2 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule2 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers