Remotery

Senior Cloud Operations Reliability Engineer – SRE

Posted Jul 21

This is a fully remote position, open to applicants in United States.

📋 Description

• The Senior Cloud Operations Reliability Engineer plays a key role in promoting operational excellence and enhancing the reliability of cloud services and supported platforms.

• This position takes ownership of essential reliability initiatives, sets up observability and service health practices, and orchestrates incident response efforts to boost service availability, resiliency, and recovery.

• Manage service reliability and operational health—create and uphold SLOs/SLIs, develop monitoring and alerting strategies, and lead enhancements that improve service availability and performance across cloud platforms.

• Direct incident response coordination and post-incident activities, including diagnosing complex production issues, performing root cause analysis, and overseeing remediation efforts with accountability for timelines and quality of resolutions.

• Design and implement automation focused on reliability, operational tooling, and runbooks to minimize manual effort, enhance response consistency, and bolster production readiness and resilience; apply Infrastructure as Code practices as needed to support recovery, reliability, and operational uniformity.

• Develop observability solutions through thorough monitoring, logging, and alerting strategies; set up event correlation and escalation protocols to guarantee swift problem detection and resolution.

• Conduct performance and capacity assessments, analyze utilization trends, identify bottlenecks, and provide recommendations for scaling, performance enhancements, capacity planning, and operational readiness of cloud services.

• Collaborate with development and engineering teams to assess deployment readiness, enhance deployment reliability, and implement operational best practices that reinforce service reliability, rollback readiness, and production supportability.

• Contribute to disaster recovery and business continuity planning, facilitate operational readiness exercises, and ensure that recovery procedures and documentation reflect the current production state and evolving business needs.

• Mentor team members and establish reliability standards and practices within Cloud Operations and supported service areas; develop and maintain operational documentation, standard operating procedures, and knowledge base resources.

• Ensure operational compliance with cloud governance, security, and compliance initiatives, including access control, tagging, logging, audit readiness, and reliability documentation.

• Carry out additional responsibilities that support the overall objectives of the role.


⛳️ Requirements

• Over 10 years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a similar field with a proven track record of managing production systems.

• Significant hands-on experience in supporting production cloud environments, specifically using Google Cloud Platform (GCP), AWS, or other equivalent cloud service providers.

• Demonstrated expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures.

• Proven experience with Infrastructure as Code tools (Terraform, Deployment Manager, CloudFormation, etc.) and best practices in version control.

• Strong foundation in incident management and post-incident review processes; experience in driving corrective actions and implementing reliability enhancements.

• Familiarity with Kubernetes operations, containerization, and orchestration platforms.

• Experience with application performance monitoring (APM) and distributed tracing.

• Experience mentoring junior engineers or leading initiatives for operational improvements.


🏝️ Benefits

• Health insurance

• 401(k) matching

• Paid time off

• Remote work options

• Professional development opportunities

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers