Director, Live Operations

Posted Sep 18

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the 24x7 operational integrity and reliability of CDM-managed production data services and workflows.

• Act as the accountable leader and escalation owner for Sev-1 and Sev-2 production incidents.

• Establish SLAs/SLOs, escalation protocols, on-call coverage models, service health metrics, and operational performance standards.

• Lead service restoration efforts, coordinate communication, and ensure follow-through on corrective actions.

• Manage CDM incident and problem management processes, encompassing severity definitions, incident command, escalation, root cause analysis, post-incident reviews, and corrective measures.

• Monitor recurring failures and leverage MTTA, MTTR, availability, incident volume, and recurrence metrics to foster systemic enhancements.

• Provide technical guidance for production services operating within AWS and related enterprise environments.

• Collaborate with Engineering on infrastructure-as-code initiatives using Terraform or similar tools, focusing on repeatable configuration, deployment, and recovery.

• Oversee operational challenges involving IAM, networking/connectivity, secure file transfers, storage, compute, logging, monitoring, and cloud dependencies.

• Champion an automation-first approach to minimize repetitive manual tasks, fragile handoffs, and reliance on specific individuals.

• Drive efforts in scripting, orchestration, automated validation, job recovery, exception handling, and self-healing methodologies when applicable.

• Partner with Data Engineering on CI/CD, APIs, ETL/data pipelines, file movement, deployment/support strategies, and production automation.

• Implement AI-assisted monitoring, troubleshooting, documentation, and workflow automation as needed.

• Develop monitoring, logging, alerting, and operational dashboards that offer actionable insights into production health.

• Set production-readiness criteria for workflows transitioning from implementation or engineering into Live Operations.

• Ensure that current runbooks, SOPs, recovery procedures, escalation paths, dependency maps, ownership, and cross-trained coverage are ready prior to production handoff.

• Clearly define operating boundaries among Live Operations, Data Engineering, Data Management, Clinical Intelligence, Product, and IT.

• Facilitate technical responses across teams during production incidents and complex operational challenges.

• Build and lead a geographically distributed technical operations team that fosters a culture of urgency, transparency, documentation, collaboration, and measurable improvement.


⛳️ Requirements

• Over 8 years of experience in production operations, cloud/platform operations, DevOps, SRE, operations engineering, data platform operations, or a similar high-availability technical environment.

• More than 4 years of experience leading technical production operations, DevOps, SRE, platform support, or operations engineering teams.

• Proven accountability for business-critical production systems operating under extended hours or 24x7 support frameworks.

• Demonstrated leadership in managing Sev-1/Sev-2 or equivalent major incidents, including incident command, restoration, root cause analysis, problem management, and corrective action follow-up.

• Strong knowledge of AWS production environments, including IAM, networking/connectivity, logging/monitoring, cloud dependencies, and operational troubleshooting.

• Proven experience with Terraform or comparable infrastructure-as-code technologies and practices for repeatable infrastructure deployment and recovery.

• Solid technical understanding of APIs, ETL/data pipelines, secure file transfers, workflow/job orchestration, SQL, automation/scripting, and production integration patterns.

• Demonstrated experience in establishing observability, monitoring, alerting, runbooks, operational dashboards, and reliability practices that can be measured.

• Proven success in automating manual production processes and reducing key-person dependencies through tools, scripting, orchestration, or platform enhancements.

• Experience managing availability, SLA/SLO attainment, MTTA, MTTR, incident volume, recurring failures, and automation coverage.

• Experience leading geographically distributed technical teams and implementing effective on-call, escalation, and coverage models.

• Capability to lead collaboration across Engineering, Product, IT, implementation, and customer-facing organizations during production incidents and initiatives focused on reliability.

• Preferred: Background in operating high-volume healthcare, financial services, SaaS, or other regulated production data environments.

• Preferred: Familiarity with MuleSoft, Flatfile, or similar enterprise integration and data ingestion platforms.

• Preferred: Experience with CI/CD tooling, cloud observability platforms, workflow orchestration, and automated recovery strategies.

• Preferred: Experience utilizing AI-assisted tools for production support, incident analysis, documentation, or operational automation.

• Preferred: Knowledge in healthcare payer data, CMS submissions, EDI/X12, enrollment, claims, risk adjustment, or related domains.


🏝️ Benefits

• Competitive salary.

• Medical, Dental, and Vision benefits.

• 401k match.

• Generous PTO plan.

People also viewed

Mercor12 hours ago

General and Operations Manager

US flagUnited States OnlyFreelanceOperations$70 – $110/hour
ApplyView job
Mission Lane16 hours ago

Senior Operations Lead, Fraud

US flagVirginia OnlyFull-timeOperations$55k – $70k/year
ApplyView job
ICF17 hours ago

Homeless Services Manager, Request and Operations

US flagCalifornia OnlyFull-timeOperations$89.6k – $152.4k/year
ApplyView job
The Cigna Group21 hours ago

Senior Process Improvement Advisor

US flagUnited States OnlyFull-timeOperations$111k – $185k/year
ApplyView job
US Foods21 hours ago

Customer Operations Specialist

US flagArizona OnlyFull-timeOperations$22 – $34/hour
ApplyView job
Mercy Health21 hours ago

Director, Ambulatory Revenue Cycle Operations

US flagOhio OnlyFull-timeOperations$120k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers