Staff Site Reliability Engineer

Posted Aug 26

This is a fully remote position, open to applicants in Florida, +3 more states.

📋 Description

• Design, implement, and sustain a highly available, scalable, and secure cloud infrastructure that supports Carrier's SaaS platforms.

• Define and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) across essential services.

• Create self-service platform capabilities for efficient and consistent service deployment and operation.

• Develop frameworks, standards, and best practices for reliability operations.

• Identify reliability bottlenecks and remove any single points of failure.

• Create automation solutions to eliminate repetitive operational tasks and minimize manual intervention.

• Design and maintain Infrastructure as Code using technologies like Terraform, AWS CloudFormation, AWS CDK, or similar.

• Build and enhance Continuous Integration/Continuous Deployment (CI/CD) pipelines.

• Implement automated remediation, self-healing features, and operational workflows.

• Design and improve observability solutions utilizing metrics, logs, traces, and distributed monitoring.

• Develop dashboards, monitoring standards, alerting frameworks, and reliability reporting systems.

• Lead responses to critical production incidents, conduct root cause analyses, and establish long-term corrective measures.

• Drive post-incident reviews that focus on systemic enhancements.

• Enhance platform performance, availability, scalability, disaster recovery, and business continuity.

• Validate backup, restoration, and recovery processes through regular testing.

• Collaborate on resilient architectures that meet reliability and compliance goals.

• Support operational readiness assessments for new services and platform capabilities.

• Mentor engineers and provide technical guidance.

• Collaborate with development teams throughout the software development lifecycle.

• Advocate for automation, ownership, continuous learning, and operational excellence.

• Stay informed about cloud infrastructure, platform engineering, AI-assisted operations, and Site Reliability Engineering (SRE) practices.

• Participate in a shared on-call rotation.

• Drive initiatives aimed at reducing operational toil through automation, self-healing systems, and platform enhancements.


⛳️ Requirements

• Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a related technical field with over 7 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Engineering, OR a Master’s degree in one of these areas with more than 5 years of relevant experience.

• At least 3 years of hands-on experience in designing and supporting large-scale cloud environments.

• A minimum of 3 years of experience with Infrastructure as Code tools, such as Terraform, CloudFormation, or AWS CDK.

• At least 3 years of experience in building and maintaining CI/CD pipelines and deployment automation.

• Over 3 years of extensive experience with Amazon Web Services, including EKS, EC2, VPC, RDS, Lambda, CloudWatch, and IAM.

• A minimum of 3 years of experience in implementing observability platforms using tools like Prometheus, Grafana, OpenTelemetry, Datadog, Splunk, or New Relic.

• Strong understanding of modern SRE principles, including reliability engineering, observability, toil reduction, incident management, and error budgets.

• Proficient scripting or programming experience with languages such as Python, Go, PowerShell, Bash, or similar.

• Strong knowledge of networking, security, systems architecture, and distributed systems concepts.

• Proven experience in leading production incident responses and conducting root cause analysis.

• Demonstrated experience with cloud-native platforms and container technologies, including Kubernetes and Docker.

• Experience operating SaaS platforms serving large-scale customer environments.

• Familiarity with supporting compliance frameworks such as SOC 2, ISO 27001, NIST, or similar standards.

• Experience in implementing AI-assisted engineering solutions to enhance operational efficiency, troubleshooting, automation, and service reliability.

• Understanding of platform engineering concepts, including Internal Developer Platforms, developer self-service capabilities, and engineering enablement practices.

• AWS certifications or other relevant cloud certifications.

• Excellent communication, leadership, and collaboration capabilities.


🏝️ Benefits

• Short-term cash incentives, contingent on plan requirements.

• Medical, Dental, Vision coverage.

• Wellness incentives.

• Retirement Benefits.

• Paid vacation days, up to 15 days.

• Paid sick days, up to 5 days.

• Paid personal leave, up to 5 days.

• Paid holidays, up to 13 days.

• Birth and adoption leave.

• Parental leave.

• Family and medical leave.

• Bereavement leave.

• Jury duty leave.

• Military leave.

• Purchased vacation options.

• Short-term and long-term disability coverage.

• Life Insurance and Accidental Death and Dismemberment coverage.

• Health Savings Account.

• Health Care Spending Account.

• Dependent Care Spending Account.

• Tuition Assistance.

People also viewed

OnePay11 hours ago

Site Reliability Engineering Lead

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$250k – $280k/year
ApplyView job
Arista Networks12 hours ago

FedRAMP Site Reliability Engineer – CloudVision

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$101k – $161k/year
ApplyView job
Octus12 hours ago

Lead DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$175k – $225k/year
ApplyView job
Tandem Diabetes Care12 hours ago

Principal Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$165k – $185k/year
ApplyView job
TechnologyAdvice15 hours ago

Senior DevOps Engineer – Contract

IN flagIndia OnlyFreelanceDevOps & Site Reliability Engineer (SRE)₹1,500 – ₹2,000/hour
ApplyView job
Bixal15 hours ago

Director of DevSecOps

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$165k – $195k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers