
Principal Site Reliability Engineer – Temp to Hire
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in United States.
• Oversee daily production support activities, encompassing intake, triage, prioritization, escalation, queue management, and change implementation.
• Develop uniform support procedures across a distributed team, focusing on shift transitions, ticket quality benchmarks, and accountability for unresolved issues.
• Direct the entire incident management process, which includes incident command, communication with stakeholders, and conducting blameless postmortems with action items tracked to resolution.
• Collaborate with Security teams to manage responses to production security incidents.
• Own the on-call strategy, including designing rotation, escalation pathways, alert tuning, and utilizing tools like PagerDuty and New Relic.
• Create runbooks that standardize procedures and facilitate first-line issue resolution.
• Define and take responsibility for SLIs and SLOs for essential services.
• Minimize MTTD and MTTR through effective instrumentation, alerting, diagnostics, and automation practices.
• Transform recurring support challenges into lasting solutions, automation, or comprehensive documentation.
• Ensure technology updates and lifecycle management for production platforms, covering cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure elements.
• Manage business continuity and disaster recovery preparedness, including backup and recovery tactics, testing recovery processes, failover capabilities, runbooks, and RTO/RPO targets.
• Lead infrastructure automation initiatives using Terraform.
• Reduce manual operational tasks and eliminate toil through automation.
• Enhance reliability measures within CI/CD pipelines, integrating automated rollback, change-risk assessments, and progressive delivery strategies.
• Ensure compliance of production systems with regulatory and compliance standards while maintaining audit readiness.
• Keep business continuity and disaster recovery documentation and evidence up to date.
• Collaborate with Security, Quality, and Compliance teams on regulatory compliance, audits, and remediation processes.
• Mentor and develop SRE and DevOps engineers through collaborative work, design and code reviews, and incident analyses.
• Advocate for documentation-first communication among distributed teams operating across multiple time zones.
• Work alongside software engineering, QA, and architecture teams to integrate reliability and operability into development practices.
• Provide insights for capacity planning and scaling strategies in collaboration with the Test team.
• Implement proactive resilience tests, such as game days.
• Align business continuity and disaster recovery strategies with application requirements and recovery goals.
• Support cloud cost management efforts, including rightsizing, reserved capacity, and governance of observability expenditures.
• Ensure all work adheres to Privacy/HIPAA and other regulatory, legal, and safety standards.
• Proven experience leading production support and incident management for live systems, including command during high-severity incidents.
• Strong knowledge of SRE principles, including SLIs/SLOs, blameless postmortems, toil reduction, and reliability engineering.
• Experience in managing on-call strategies, including rotation design, alert tuning, and escalation procedures.
• Proficiency in Terraform or similar Infrastructure as Code (IaC) solutions at scale, including module design, state management, and policy-as-code guidelines.
• Hands-on experience in constructing CI/CD pipelines with reliability safeguards using tools such as GitHub Actions, Octopus Deploy, or Azure DevOps.
• Extensive experience with at least one leading cloud platform: AWS, Azure, or GCP.
• Familiarity with Docker and Kubernetes.
• Understanding of observability tools such as Prometheus, Grafana, Datadog, CloudWatch, or ELK/OpenSearch.
• Experience in designing and testing disaster recovery plans, including backup/restore, failover, and RTO/RPO validation.
• Knowledge of cloud security and compliance practices, covering IAM, network segmentation, encryption, vulnerability and patch management, and cost optimization.
• Proficient in at least one scripting or programming language, such as Python, Go, or Bash.
• 10+ years of experience in Site Reliability Engineering, DevOps, or infrastructure engineering.
• 2+ years of experience mentoring or technically leading other engineers, including remote, offshore, or contracted staff.
• Experience in FDA and ISO regulated sectors and agile methodologies is preferred.
• Bachelor’s degree in Computer Science or an equivalent combination of education and relevant job experience, including technical training and certifications; demonstrated production experience is prioritized over a degree.
• Relevant cloud certifications are preferred.
• Must be eligible to work for any employer in the U.S.; the employer cannot sponsor or assume sponsorship of an employment visa.
• Equipment necessary for the role will be provided.
• Training will be conducted virtually.
• Competitive compensation package that includes bonuses.
• Comprehensive benefits package.
• Tandem-sponsored benefits contingent upon transitioning from temporary to regular full-time status.
• Commitment to equal opportunity and an inclusive work environment.
• A positive workplace that celebrates achievements and supports employee well-being.
HubSpot
InfluxData
LocalStack
Amigo Tech
Get handpicked remote jobs straight to your inbox weekly.