Remotery

Senior Site Reliability Engineer – Infra Ops

Posted Aug 5

This is a fully remote position, open to applicants in California.

📋 Description

• Design, construct, and manage secure, highly available, and scalable Kubernetes platforms in hybrid and public cloud settings.

• Create reusable infrastructure-as-code modules and automated delivery workflows utilizing Terraform.

• Develop backend services, internal tools, and operational automation using Go, Python, or JavaScript/TypeScript.

• Collaborate with engineering and product teams to translate workload requirements into dependable, secure, efficient, and cost-effective technical designs.

• Enhance the production lifecycle through CI/CD, deployment automation, progressive delivery, and operational ownership.

• Establish observability practices encompassing metrics, logs, traces, alerting, and dashboards.

• Engage in on-call duties, lead incident response efforts, perform root-cause analysis, and facilitate blameless postmortems and corrective measures.

• Set reliability targets utilizing SLIs, SLOs, error budgets, capacity planning, disaster-recovery testing, and resilience enhancements.

• Integrate security and compliance into platform operations and collaborate with the Security team.

• Utilize AI-assisted and data-driven operational strategies to enhance detection, minimize alert noise, expedite root-cause analysis, and automate workflows.

• Carry out code reviews, documentation, knowledge sharing, and mentorship.

• Guide and support team development.


⛳️ Requirements

• Over 5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or a closely related software engineering role supporting production systems.

• Extensive hands-on experience with Kubernetes, including the design, operation, security, and troubleshooting of production clusters and containerized workloads at scale.

• Proficient in Terraform, with experience in creating reusable modules, managing state and environments, and automated infrastructure delivery workflows.

• Practical software development experience in Go, Python, or JavaScript/TypeScript.

• Demonstrated experience in enhancing the reliability, performance, scalability, or cost-effectiveness of distributed systems in production.

• Familiarity with cloud infrastructure and networking concepts, including IAM, DNS, load balancing, routing, service networking, and secure connectivity.

• Strong skills in observability and troubleshooting using metrics, logs, traces, alerting, and incident data.

• Experience working with SLIs, SLOs, error budgets, incident management, postmortems, and disaster recovery protocols.

• Knowledge of CI/CD, GitOps or deployment automation, and canary or blue-green deployment strategies.

• Security-focused infrastructure approach with experience collaborating with Security and engineering teams in regulated or high-availability environments.

• Excellent written and verbal communication skills, with strong ownership and judgment in balancing speed, risk, and operational excellence.

• Familiarity with AI-assisted tooling is a plus.


🏝️ Benefits

• Flexible work environment.

• Inclusive work environment.

• Equal opportunity employer.

• Interview accommodations or assistance for candidates with disabilities.

• Compensation package including base pay.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers