Head of Cloud Operations

atHavocAIRemoteUS flagUnited StatesFull-timeOperationsLead$200k – $220k/year

Posted 3 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Lead and oversee the SRE and DevOps teams, establishing both technical and operational strategies focused on reliability, automation, and safe delivery.

• Recruit, mentor, and manage engineers across both teams while assessing performance.

• Take ownership of reliability and delivery standards, which include SLIs, SLOs, error budgets, on-call health, CI/CD, and infrastructure automation.

• Define clear ownership and operational expectations for cloud reliability and delivery.

• Collaborate with engineering leaders to identify systemic reliability challenges and prioritize enhancements.

• Define and manage change classifications, approval processes, impact assessments, rollback strategies, and documentation of approvals.

• Integrate change management with GitOps workflows utilizing merged, signed, peer-reviewed pull requests.

• Manage maintenance windows, freeze periods, emergency change procedures, release schedules, versioning strategies, and promotions across various environments and tenants.

• Set pre-deployment verification standards and coordinate releases with Cloud Platform, Backend, Autonomy, and Frontend teams.

• Maintain audit-ready documentation for deployments and releases.

• Oversee the incident management framework, which includes declarations, severity classifications, escalation paths, incident command responsibilities, communications, and notifications.

• Lead incident command during significant events and facilitate blameless post-incident reviews.

• Handle government notification requirements for incidents impacting authorized systems.

• Differentiate between incidents and problems and drive the analysis of recurring issues.

• Keep authoritative records of deployed systems, versions, digests, dependencies, baselines, and configuration changes.

• Provide audit-ready documentation for the ISSO and monitor POA&M remediation commitments.

• Convert security and compliance controls into actionable engineering processes.

• Develop automation to minimize manual compliance tasks and enhance the reliability of operational evidence.

• Implement effective change, release, incident, reliability, compliance, and remediation processes within the first year.


⛳️ Requirements

• Over 8 years of relevant experience in change management, release management, incident management, technical program management, service management, SRE, DevOps, platform engineering, or related fields.

• Proven experience in leading or managing SRE, DevOps, or Platform Engineering teams, including hiring, coaching, and evaluating performance.

• Deep understanding of modern cloud operations, software delivery, infrastructure automation, and production reliability.

• Experience in establishing and managing change, release, and incident management processes in complex technical settings.

• Demonstrated ability to coordinate complex initiatives across engineering teams and with stakeholders not directly managed.

• Technical proficiency to evaluate and challenge engineering impact assessments, deployment strategies, rollback plans, and root-cause analyses.

• Ability to stay composed, decisive, and directive during active incidents, making sound decisions under pressure.

• Excellent written communication skills, particularly for incident communications, post-incident reports, operational documentation, and executive updates.

• Capability to develop scalable processes that ensure appropriate controls without hindering engineering teams unnecessarily.

• Strong sense of ownership, judgment, and comfort operating in a fast-paced and ambiguous environment.

• Must be a U.S. Citizen and eligible to obtain and maintain a U.S. Government security clearance.

• Previous incident command experience in a regulated, defense, government, or safety-critical environment.

• Experience managing cloud systems that are subject to U.S. Government authorization or compliance standards.

• Familiarity with POA&Ms, security authorization processes, configuration baselines, and audit evidence management.

• Experience in implementing GitOps-based change and release methodologies.

• Knowledge of ITIL practices or equivalent hands-on experience in developing effective service management processes.

• Proficiency with tools such as Jira, PagerDuty, status pages, runbook platforms, and incident management systems.

• Experience with Kubernetes, infrastructure as code, CI/CD platforms, observability systems, and modern cloud infrastructure.


🏝️ Benefits

• 100% Employer-covered Health, Dental, and Vision Insurance for you and your family.

• Life Insurance (Employer Paid).

• Opportunity to participate in the company's 401k program (with matching).

• Unlimited PTO policy with a mandatory minimum of 2 weeks.

• Equity Package.

• Work/Home Office Stipend.

• Global Entry.

• 16 Weeks of Paid Parental Leave.

• Monthly Health and Wellness Stipend.

• Bonus.

People also viewed

JUFA1 day ago

Hotel Operations Coordinator

AT flagAustria OnlyFull-timeOperations€3,000/month
ApplyView job
Astro Pak1 day ago

Operations Buyer

US flagUnited States OnlyFull-timeOperations
ApplyView job
CareMetx, LLC1 day ago

Operations Manager

US flagUnited States OnlyFull-timeOperations$70k – $80k/year
ApplyView job
Zocdoc1 day ago

Senior Manager, Talent Operations

US flagCalifornia, +1 more stateFull-timeOperations$159k – $215k/year
ApplyView job
Delinea1 day ago

Principal, Major Event Strategy & Operations

US flagUnited States OnlyFull-timeOperations$156k – $195k/year
ApplyView job
Quantanite1 day ago

Associate – Operations

BD flagBangladesh OnlyFreelanceOperations
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers