Remotery

Principal Site Reliability Engineer – Temp to Hire

Posted 15 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Oversee daily production support activities, encompassing intake, triage, prioritization, escalation, queue management, and change implementation.

• Develop uniform support procedures across a distributed team, focusing on shift transitions, ticket quality benchmarks, and accountability for unresolved issues.

• Direct the entire incident management process, which includes incident command, communication with stakeholders, and conducting blameless postmortems with action items tracked to resolution.

• Collaborate with Security teams to manage responses to production security incidents.

• Own the on-call strategy, including designing rotation, escalation pathways, alert tuning, and utilizing tools like PagerDuty and New Relic.

• Create runbooks that standardize procedures and facilitate first-line issue resolution.

• Define and take responsibility for SLIs and SLOs for essential services.

• Minimize MTTD and MTTR through effective instrumentation, alerting, diagnostics, and automation practices.

• Transform recurring support challenges into lasting solutions, automation, or comprehensive documentation.

• Ensure technology updates and lifecycle management for production platforms, covering cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure elements.

• Manage business continuity and disaster recovery preparedness, including backup and recovery tactics, testing recovery processes, failover capabilities, runbooks, and RTO/RPO targets.

• Lead infrastructure automation initiatives using Terraform.

• Reduce manual operational tasks and eliminate toil through automation.

• Enhance reliability measures within CI/CD pipelines, integrating automated rollback, change-risk assessments, and progressive delivery strategies.

• Ensure compliance of production systems with regulatory and compliance standards while maintaining audit readiness.

• Keep business continuity and disaster recovery documentation and evidence up to date.

• Collaborate with Security, Quality, and Compliance teams on regulatory compliance, audits, and remediation processes.

• Mentor and develop SRE and DevOps engineers through collaborative work, design and code reviews, and incident analyses.

• Advocate for documentation-first communication among distributed teams operating across multiple time zones.

• Work alongside software engineering, QA, and architecture teams to integrate reliability and operability into development practices.

• Provide insights for capacity planning and scaling strategies in collaboration with the Test team.

• Implement proactive resilience tests, such as game days.

• Align business continuity and disaster recovery strategies with application requirements and recovery goals.

• Support cloud cost management efforts, including rightsizing, reserved capacity, and governance of observability expenditures.

• Ensure all work adheres to Privacy/HIPAA and other regulatory, legal, and safety standards.


⛳️ Requirements

• Proven experience leading production support and incident management for live systems, including command during high-severity incidents.

• Strong knowledge of SRE principles, including SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering.

• Experience in managing on-call strategies, including rotation design, alert tuning, and escalation procedures.

• Proficiency in Terraform or similar Infrastructure as Code (IaC) solutions at scale, including module design, state management, and policy-as-code guidelines.

• Hands-on experience in constructing CI/CD pipelines with reliability safeguards using tools such as GitHub Actions, Octopus Deploy, or Azure DevOps.

• Extensive experience with at least one leading cloud platform: AWS, Azure, or GCP.

• Familiarity with Docker and Kubernetes.

• Understanding of observability tools such as Prometheus, Grafana, Datadog, CloudWatch, or ELK/OpenSearch.

• Experience in designing and testing disaster recovery plans, including backup/restore, failover, and RTO/RPO validation.

• Knowledge of cloud security and compliance practices, covering IAM, network segmentation, encryption, vulnerability and patch management, and cost optimization.

• Proficient in at least one scripting or programming language, such as Python, Go, or Bash.

• 10+ years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering.

• 2+ years of experience mentoring or technically leading other engineers, including remote, offshore, or contracted staff.

• Experience in FDA and ISO regulated sectors and agile methodologies is preferred.

• Bachelor’s degree in Computer Science or an equivalent combination of education and relevant job experience, including technical training and certifications; demonstrated production experience is prioritized over a degree.

• Relevant cloud certifications are preferred.

• Must be eligible to work for any employer in the U.S.; the employer cannot sponsor or assume sponsorship of an employment visa.


🏝️ Benefits

• Equipment necessary for the role will be provided.

• Training will be conducted virtually.

• Competitive compensation package that includes bonuses.

• Comprehensive benefits package.

• Tandem-sponsored benefits contingent upon transitioning from temporary to regular full-time status.

• Commitment to equal opportunity and an inclusive work environment.

• A positive workplace that celebrates achievements and supports employee well-being.

People also viewed

HubSpot13 hours ago

Principal Software Engineer, Developer Acceleration – Release Engineering

IE flagIreland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$242k – $387.2k/year
ApplyView job
InfluxData13 hours ago

DevOps Engineer

GB flagUnited Kingdom, +7 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
LocalStack14 hours ago

Senior DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€50k – €78k/year
ApplyView job
Amigo Tech15 hours ago

DevOps Engineer, Mid-Level

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Penn Interactive15 hours ago

Senior Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$145k – $193k/year
ApplyView job
Smarthis15 hours ago

Senior Cloud Platform – DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers