
Staff Site Reliability Engineer
Posted Aug 26

Posted Aug 26
This is a fully remote position, open to applicants in Florida, +3 more states.
• Design, implement, and sustain a highly available, scalable, and secure cloud infrastructure that supports Carrier's SaaS platforms.
• Define and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs) across essential services.
• Create self-service platform capabilities for efficient and consistent service deployment and operation.
• Develop frameworks, standards, and best practices for reliability operations.
• Identify reliability bottlenecks and remove any single points of failure.
• Create automation solutions to eliminate repetitive operational tasks and minimize manual intervention.
• Design and maintain Infrastructure as Code using technologies like Terraform, AWS CloudFormation, AWS CDK, or similar.
• Build and enhance Continuous Integration/Continuous Deployment (CI/CD) pipelines.
• Implement automated remediation, self-healing features, and operational workflows.
• Design and improve observability solutions utilizing metrics, logs, traces, and distributed monitoring.
• Develop dashboards, monitoring standards, alerting frameworks, and reliability reporting systems.
• Lead responses to critical production incidents, conduct root cause analyses, and establish long-term corrective measures.
• Drive post-incident reviews that focus on systemic enhancements.
• Enhance platform performance, availability, scalability, disaster recovery, and business continuity.
• Validate backup, restoration, and recovery processes through regular testing.
• Collaborate on resilient architectures that meet reliability and compliance goals.
• Support operational readiness assessments for new services and platform capabilities.
• Mentor engineers and provide technical guidance.
• Collaborate with development teams throughout the software development lifecycle.
• Advocate for automation, ownership, continuous learning, and operational excellence.
• Stay informed about cloud infrastructure, platform engineering, AI-assisted operations, and Site Reliability Engineering (SRE) practices.
• Participate in a shared on-call rotation.
• Drive initiatives aimed at reducing operational toil through automation, self-healing systems, and platform enhancements.
• Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a related technical field with over 7 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Engineering, OR a Master’s degree in one of these areas with more than 5 years of relevant experience.
• At least 3 years of hands-on experience in designing and supporting large-scale cloud environments.
• A minimum of 3 years of experience with Infrastructure as Code tools, such as Terraform, CloudFormation, or AWS CDK.
• At least 3 years of experience in building and maintaining CI/CD pipelines and deployment automation.
• Over 3 years of extensive experience with Amazon Web Services, including EKS, EC2, VPC, RDS, Lambda, CloudWatch, and IAM.
• A minimum of 3 years of experience in implementing observability platforms using tools like Prometheus, Grafana, OpenTelemetry, Datadog, Splunk, or New Relic.
• Strong understanding of modern SRE principles, including reliability engineering, observability, toil reduction, incident management, and error budgets.
• Proficient scripting or programming experience with languages such as Python, Go, PowerShell, Bash, or similar.
• Strong knowledge of networking, security, systems architecture, and distributed systems concepts.
• Proven experience in leading production incident responses and conducting root cause analysis.
• Demonstrated experience with cloud-native platforms and container technologies, including Kubernetes and Docker.
• Experience operating SaaS platforms serving large-scale customer environments.
• Familiarity with supporting compliance frameworks such as SOC 2, ISO 27001, NIST, or similar standards.
• Experience in implementing AI-assisted engineering solutions to enhance operational efficiency, troubleshooting, automation, and service reliability.
• Understanding of platform engineering concepts, including Internal Developer Platforms, developer self-service capabilities, and engineering enablement practices.
• AWS certifications or other relevant cloud certifications.
• Excellent communication, leadership, and collaboration capabilities.
• Short-term cash incentives, contingent on plan requirements.
• Medical, Dental, Vision coverage.
• Wellness incentives.
• Retirement Benefits.
• Paid vacation days, up to 15 days.
• Paid sick days, up to 5 days.
• Paid personal leave, up to 5 days.
• Paid holidays, up to 13 days.
• Birth and adoption leave.
• Parental leave.
• Family and medical leave.
• Bereavement leave.
• Jury duty leave.
• Military leave.
• Purchased vacation options.
• Short-term and long-term disability coverage.
• Life Insurance and Accidental Death and Dismemberment coverage.
• Health Savings Account.
• Health Care Spending Account.
• Dependent Care Spending Account.
• Tuition Assistance.
OnePay
Arista Networks
Octus
Tandem Diabetes Care
Get handpicked remote jobs straight to your inbox weekly.