Principal Site Reliability Engineer

Posted 10 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Oversee daily production support activities, including intake, triage, prioritization, escalation, queue management, and change execution.

• Implement standardized support practices across a distributed team, encompassing shift handoffs, ticket quality standards, and accountability for unresolved issues.

• Manage incident response from start to finish, including incident command, stakeholder communication, and conducting blameless postmortems with corrective actions monitored until completion.

• Collaborate with Security teams to coordinate responses to production security incidents.

• Develop and manage on-call strategies, rotation structures, escalation protocols, alert tuning, and tools such as PagerDuty and New Relic.

• Act as a senior escalation point for critical incidents.

• Create runbooks for common failure scenarios and first-response solutions.

• Define and oversee SLIs and SLOs for essential services.

• Minimize MTTD and MTTR through effective instrumentation, alerting, diagnostics, and automation.

• Transform ongoing support challenges into permanent solutions, automation, or documentation.

• Keep abreast of technology trends and manage the lifecycle of cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components.

• Ensure readiness for business continuity and disaster recovery, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and RTO/RPO objectives.

• Drive infrastructure automation using Terraform.

• Reduce manual labor through automation.

• Implement reliability measures in CI/CD pipelines, including automated rollback, change-risk assessments, and progressive delivery.

• Maintain production systems in compliance with regulatory and compliance requirements.

• Update business continuity and disaster recovery documentation and provide audit evidence.

• Collaborate with Security, Quality, and Compliance teams on audits, compliance, and remediation efforts.

• Mentor SRE and DevOps engineers through pair programming, design and code reviews, and incident debriefs.

• Promote open communication among junior and contract engineers.

• Emphasize documentation-first practices across distributed teams in different time zones.

• Work alongside software engineering, QA, and architecture teams to integrate reliability into the development lifecycle.

• Assist the Test team in informing capacity planning and scaling strategies.

• Introduce proactive resilience testing, such as game days.

• Align business continuity and disaster recovery capabilities with application requirements.

• Support cloud cost optimization through rightsizing, reserved capacity, and observability spending governance.

• Ensure all work adheres to company policies and applicable Privacy/HIPAA, regulatory, legal, and safety standards.


⛳️ Requirements

• Proven experience in leading production support and incident management for production systems, including incident command for high-severity events.

• Strong foundation in SRE principles, such as SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering.

• Demonstrated experience in managing on-call strategies, including rotation design, alert tuning, and escalation procedures.

• Proficient in Terraform or similar Infrastructure as Code (IaC) tools at scale, including module design, state management, and policy-as-code practices.

• Hands-on experience constructing CI/CD pipelines with reliability measures using GitHub Actions, Octopus Deploy, or Azure DevOps.

• Extensive experience with at least one leading cloud platform: AWS, Azure, or GCP.

• Familiarity with Docker and Kubernetes.

• Working knowledge of observability tools such as Prometheus, Grafana, Datadog, CloudWatch, and ELK/OpenSearch.

• Experience in designing and testing disaster recovery plans, including backup/restore, failover processes, and validating RTO/RPO metrics.

• Understanding of cloud security and compliance practices, including IAM, network segmentation, encryption, vulnerability management, and cloud cost optimization.

• Proficient in at least one scripting or programming language such as Python, Go, or Bash.

• Experience in FDA and ISO regulated industries and familiar with agile methodologies is preferred.

• B.S. in Computer Science or equivalent education and applicable job experience, including technical training and certifications in networks, servers, and cloud infrastructure; demonstrated production experience is valued more than formal education.

• Relevant cloud certifications such as AWS/Azure/GCP Professional or Architect level are preferred.

• 10+ years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering.

• 2+ years of experience mentoring or technically leading other engineers, including remote, offshore, or contracted partner engineers.

• Must reside within the United States.

• Successful completion of pre-employment drug screening and background checks.

• Adherence to applicable company policies and Privacy/HIPAA, regulatory, legal, and safety requirements.


🏝️ Benefits

• Medical, dental, and vision benefits available from day one.

• Health savings accounts offered.

• Flexible savings accounts available.

• 11 paid holidays each year.

• Minimum of 20 days of paid time off, with accrual beginning on day one.

• 401(k) plan with company matching.

• Employee Stock Purchase plan accessible.

• Equipment provided for work purposes.

• Virtual training opportunities available.

• Competitive compensation package with bonuses.

• Joy-focused workplace that promotes well-being, achievement, growth, fun, and camaraderie.

People also viewed

OnePay8 hours ago

Site Reliability Engineering Lead

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$250k – $280k/year
ApplyView job
Arista Networks9 hours ago

FedRAMP Site Reliability Engineer – CloudVision

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$101k – $161k/year
ApplyView job
Octus9 hours ago

Lead DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$175k – $225k/year
ApplyView job
TechnologyAdvice12 hours ago

Senior DevOps Engineer – Contract

IN flagIndia OnlyFreelanceDevOps & Site Reliability Engineer (SRE)₹1,500 – ₹2,000/hour
ApplyView job
Bixal12 hours ago

Director of DevSecOps

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$165k – $195k/year
ApplyView job
Sphera13 hours ago

Cybersecurity Engineer – DevOps

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$112k – $178k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers