
Head of Cloud Operations
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in United States.
• Lead and oversee the SRE and DevOps teams, establishing both technical and operational strategies focused on reliability, automation, and safe delivery.
• Recruit, mentor, and manage engineers across both teams while assessing performance.
• Take ownership of reliability and delivery standards, which include SLIs, SLOs, error budgets, on-call health, CI/CD, and infrastructure automation.
• Define clear ownership and operational expectations for cloud reliability and delivery.
• Collaborate with engineering leaders to identify systemic reliability challenges and prioritize enhancements.
• Define and manage change classifications, approval processes, impact assessments, rollback strategies, and documentation of approvals.
• Integrate change management with GitOps workflows utilizing merged, signed, peer-reviewed pull requests.
• Manage maintenance windows, freeze periods, emergency change procedures, release schedules, versioning strategies, and promotions across various environments and tenants.
• Set pre-deployment verification standards and coordinate releases with Cloud Platform, Backend, Autonomy, and Frontend teams.
• Maintain audit-ready documentation for deployments and releases.
• Oversee the incident management framework, which includes declarations, severity classifications, escalation paths, incident command responsibilities, communications, and notifications.
• Lead incident command during significant events and facilitate blameless post-incident reviews.
• Handle government notification requirements for incidents impacting authorized systems.
• Differentiate between incidents and problems and drive the analysis of recurring issues.
• Keep authoritative records of deployed systems, versions, digests, dependencies, baselines, and configuration changes.
• Provide audit-ready documentation for the ISSO and monitor POA&M remediation commitments.
• Convert security and compliance controls into actionable engineering processes.
• Develop automation to minimize manual compliance tasks and enhance the reliability of operational evidence.
• Implement effective change, release, incident, reliability, compliance, and remediation processes within the first year.
• Over 8 years of relevant experience in change management, release management, incident management, technical program management, service management, SRE, DevOps, platform engineering, or related fields.
• Proven experience in leading or managing SRE, DevOps, or Platform Engineering teams, including hiring, coaching, and evaluating performance.
• Deep understanding of modern cloud operations, software delivery, infrastructure automation, and production reliability.
• Experience in establishing and managing change, release, and incident management processes in complex technical settings.
• Demonstrated ability to coordinate complex initiatives across engineering teams and with stakeholders not directly managed.
• Technical proficiency to evaluate and challenge engineering impact assessments, deployment strategies, rollback plans, and root-cause analyses.
• Ability to stay composed, decisive, and directive during active incidents, making sound decisions under pressure.
• Excellent written communication skills, particularly for incident communications, post-incident reports, operational documentation, and executive updates.
• Capability to develop scalable processes that ensure appropriate controls without hindering engineering teams unnecessarily.
• Strong sense of ownership, judgment, and comfort operating in a fast-paced and ambiguous environment.
• Must be a U.S. Citizen and eligible to obtain and maintain a U.S. Government security clearance.
• Previous incident command experience in a regulated, defense, government, or safety-critical environment.
• Experience managing cloud systems that are subject to U.S. Government authorization or compliance standards.
• Familiarity with POA&Ms, security authorization processes, configuration baselines, and audit evidence management.
• Experience in implementing GitOps-based change and release methodologies.
• Knowledge of ITIL practices or equivalent hands-on experience in developing effective service management processes.
• Proficiency with tools such as Jira, PagerDuty, status pages, runbook platforms, and incident management systems.
• Experience with Kubernetes, infrastructure as code, CI/CD platforms, observability systems, and modern cloud infrastructure.
• 100% Employer-covered Health, Dental, and Vision Insurance for you and your family.
• Life Insurance (Employer Paid).
• Opportunity to participate in the company's 401k program (with matching).
• Unlimited PTO policy with a mandatory minimum of 2 weeks.
• Equity Package.
• Work/Home Office Stipend.
• Global Entry.
• 16 Weeks of Paid Parental Leave.
• Monthly Health and Wellness Stipend.
• Bonus.
CareMetx, LLC
Zocdoc
Get handpicked remote jobs straight to your inbox weekly.