
Senior Cloud Operations Reliability Engineer – SRE
Posted Jul 21

Posted Jul 21
This is a fully remote position, open to applicants in United States.
• The Senior Cloud Operations Reliability Engineer plays a key role in promoting operational excellence and enhancing the reliability of cloud services and supported platforms.
• This position takes ownership of essential reliability initiatives, sets up observability and service health practices, and orchestrates incident response efforts to boost service availability, resiliency, and recovery.
• Manage service reliability and operational health—create and uphold SLOs/SLIs, develop monitoring and alerting strategies, and lead enhancements that improve service availability and performance across cloud platforms.
• Direct incident response coordination and post-incident activities, including diagnosing complex production issues, performing root cause analysis, and overseeing remediation efforts with accountability for timelines and quality of resolutions.
• Design and implement automation focused on reliability, operational tooling, and runbooks to minimize manual effort, enhance response consistency, and bolster production readiness and resilience; apply Infrastructure as Code practices as needed to support recovery, reliability, and operational uniformity.
• Develop observability solutions through thorough monitoring, logging, and alerting strategies; set up event correlation and escalation protocols to guarantee swift problem detection and resolution.
• Conduct performance and capacity assessments, analyze utilization trends, identify bottlenecks, and provide recommendations for scaling, performance enhancements, capacity planning, and operational readiness of cloud services.
• Collaborate with development and engineering teams to assess deployment readiness, enhance deployment reliability, and implement operational best practices that reinforce service reliability, rollback readiness, and production supportability.
• Contribute to disaster recovery and business continuity planning, facilitate operational readiness exercises, and ensure that recovery procedures and documentation reflect the current production state and evolving business needs.
• Mentor team members and establish reliability standards and practices within Cloud Operations and supported service areas; develop and maintain operational documentation, standard operating procedures, and knowledge base resources.
• Ensure operational compliance with cloud governance, security, and compliance initiatives, including access control, tagging, logging, audit readiness, and reliability documentation.
• Carry out additional responsibilities that support the overall objectives of the role.
• Over 10 years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a similar field with a proven track record of managing production systems.
• Significant hands-on experience in supporting production cloud environments, specifically using Google Cloud Platform (GCP), AWS, or other equivalent cloud service providers.
• Demonstrated expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures.
• Proven experience with Infrastructure as Code tools (Terraform, Deployment Manager, CloudFormation, etc.) and best practices in version control.
• Strong foundation in incident management and post-incident review processes; experience in driving corrective actions and implementing reliability enhancements.
• Familiarity with Kubernetes operations, containerization, and orchestration platforms.
• Experience with application performance monitoring (APM) and distributed tracing.
• Experience mentoring junior engineers or leading initiatives for operational improvements.
• Health insurance
• 401(k) matching
• Paid time off
• Remote work options
• Professional development opportunities
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.