
Principal Site Reliability Engineer
Posted 10 hours ago

Posted 10 hours ago
This is a fully remote position, open to applicants in United States.
• Oversee daily production support activities, including intake, triage, prioritization, escalation, queue management, and change execution.
• Implement standardized support practices across a distributed team, encompassing shift handoffs, ticket quality standards, and accountability for unresolved issues.
• Manage incident response from start to finish, including incident command, stakeholder communication, and conducting blameless postmortems with corrective actions monitored until completion.
• Collaborate with Security teams to coordinate responses to production security incidents.
• Develop and manage on-call strategies, rotation structures, escalation protocols, alert tuning, and tools such as PagerDuty and New Relic.
• Act as a senior escalation point for critical incidents.
• Create runbooks for common failure scenarios and first-response solutions.
• Define and oversee SLIs and SLOs for essential services.
• Minimize MTTD and MTTR through effective instrumentation, alerting, diagnostics, and automation.
• Transform ongoing support challenges into permanent solutions, automation, or documentation.
• Keep abreast of technology trends and manage the lifecycle of cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components.
• Ensure readiness for business continuity and disaster recovery, including backup and recovery strategies, recovery testing, failover capabilities, recovery runbooks, and RTO/RPO objectives.
• Drive infrastructure automation using Terraform.
• Reduce manual labor through automation.
• Implement reliability measures in CI/CD pipelines, including automated rollback, change-risk assessments, and progressive delivery.
• Maintain production systems in compliance with regulatory and compliance requirements.
• Update business continuity and disaster recovery documentation and provide audit evidence.
• Collaborate with Security, Quality, and Compliance teams on audits, compliance, and remediation efforts.
• Mentor SRE and DevOps engineers through pair programming, design and code reviews, and incident debriefs.
• Promote open communication among junior and contract engineers.
• Emphasize documentation-first practices across distributed teams in different time zones.
• Work alongside software engineering, QA, and architecture teams to integrate reliability into the development lifecycle.
• Assist the Test team in informing capacity planning and scaling strategies.
• Introduce proactive resilience testing, such as game days.
• Align business continuity and disaster recovery capabilities with application requirements.
• Support cloud cost optimization through rightsizing, reserved capacity, and observability spending governance.
• Ensure all work adheres to company policies and applicable Privacy/HIPAA, regulatory, legal, and safety standards.
• Proven experience in leading production support and incident management for production systems, including incident command for high-severity events.
• Strong foundation in SRE principles, such as SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering.
• Demonstrated experience in managing on-call strategies, including rotation design, alert tuning, and escalation procedures.
• Proficient in Terraform or similar Infrastructure as Code (IaC) tools at scale, including module design, state management, and policy-as-code practices.
• Hands-on experience constructing CI/CD pipelines with reliability measures using GitHub Actions, Octopus Deploy, or Azure DevOps.
• Extensive experience with at least one leading cloud platform: AWS, Azure, or GCP.
• Familiarity with Docker and Kubernetes.
• Working knowledge of observability tools such as Prometheus, Grafana, Datadog, CloudWatch, and ELK/OpenSearch.
• Experience in designing and testing disaster recovery plans, including backup/restore, failover processes, and validating RTO/RPO metrics.
• Understanding of cloud security and compliance practices, including IAM, network segmentation, encryption, vulnerability management, and cloud cost optimization.
• Proficient in at least one scripting or programming language such as Python, Go, or Bash.
• Experience in FDA and ISO regulated industries and familiar with agile methodologies is preferred.
• B.S. in Computer Science or equivalent education and applicable job experience, including technical training and certifications in networks, servers, and cloud infrastructure; demonstrated production experience is valued more than formal education.
• Relevant cloud certifications such as AWS/Azure/GCP Professional or Architect level are preferred.
• 10+ years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering.
• 2+ years of experience mentoring or technically leading other engineers, including remote, offshore, or contracted partner engineers.
• Must reside within the United States.
• Successful completion of pre-employment drug screening and background checks.
• Adherence to applicable company policies and Privacy/HIPAA, regulatory, legal, and safety requirements.
• Medical, dental, and vision benefits available from day one.
• Health savings accounts offered.
• Flexible savings accounts available.
• 11 paid holidays each year.
• Minimum of 20 days of paid time off, with accrual beginning on day one.
• 401(k) plan with company matching.
• Employee Stock Purchase plan accessible.
• Equipment provided for work purposes.
• Virtual training opportunities available.
• Competitive compensation package with bonuses.
• Joy-focused workplace that promotes well-being, achievement, growth, fun, and camaraderie.
OnePay
Arista Networks
Octus
TechnologyAdvice
Get handpicked remote jobs straight to your inbox weekly.