
Director, Live Operations
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in United States.
• Take ownership of the 24x7 operational integrity and reliability of CDM-managed production data services and workflows.
• Act as the accountable leader and escalation owner for Sev-1 and Sev-2 production incidents.
• Establish SLAs/SLOs, escalation protocols, on-call coverage models, service health metrics, and operational performance standards.
• Lead service restoration efforts, coordinate communication, and ensure follow-through on corrective actions.
• Manage CDM incident and problem management processes, encompassing severity definitions, incident command, escalation, root cause analysis, post-incident reviews, and corrective measures.
• Monitor recurring failures and leverage MTTA, MTTR, availability, incident volume, and recurrence metrics to foster systemic enhancements.
• Provide technical guidance for production services operating within AWS and related enterprise environments.
• Collaborate with Engineering on infrastructure-as-code initiatives using Terraform or similar tools, focusing on repeatable configuration, deployment, and recovery.
• Oversee operational challenges involving IAM, networking/connectivity, secure file transfers, storage, compute, logging, monitoring, and cloud dependencies.
• Champion an automation-first approach to minimize repetitive manual tasks, fragile handoffs, and reliance on specific individuals.
• Drive efforts in scripting, orchestration, automated validation, job recovery, exception handling, and self-healing methodologies when applicable.
• Partner with Data Engineering on CI/CD, APIs, ETL/data pipelines, file movement, deployment/support strategies, and production automation.
• Implement AI-assisted monitoring, troubleshooting, documentation, and workflow automation as needed.
• Develop monitoring, logging, alerting, and operational dashboards that offer actionable insights into production health.
• Set production-readiness criteria for workflows transitioning from implementation or engineering into Live Operations.
• Ensure that current runbooks, SOPs, recovery procedures, escalation paths, dependency maps, ownership, and cross-trained coverage are ready prior to production handoff.
• Clearly define operating boundaries among Live Operations, Data Engineering, Data Management, Clinical Intelligence, Product, and IT.
• Facilitate technical responses across teams during production incidents and complex operational challenges.
• Build and lead a geographically distributed technical operations team that fosters a culture of urgency, transparency, documentation, collaboration, and measurable improvement.
• Over 8 years of experience in production operations, cloud/platform operations, DevOps, SRE, operations engineering, data platform operations, or a similar high-availability technical environment.
• More than 4 years of experience leading technical production operations, DevOps, SRE, platform support, or operations engineering teams.
• Proven accountability for business-critical production systems operating under extended hours or 24x7 support frameworks.
• Demonstrated leadership in managing Sev-1/Sev-2 or equivalent major incidents, including incident command, restoration, root cause analysis, problem management, and corrective action follow-up.
• Strong knowledge of AWS production environments, including IAM, networking/connectivity, logging/monitoring, cloud dependencies, and operational troubleshooting.
• Proven experience with Terraform or comparable infrastructure-as-code technologies and practices for repeatable infrastructure deployment and recovery.
• Solid technical understanding of APIs, ETL/data pipelines, secure file transfers, workflow/job orchestration, SQL, automation/scripting, and production integration patterns.
• Demonstrated experience in establishing observability, monitoring, alerting, runbooks, operational dashboards, and reliability practices that can be measured.
• Proven success in automating manual production processes and reducing key-person dependencies through tools, scripting, orchestration, or platform enhancements.
• Experience managing availability, SLA/SLO attainment, MTTA, MTTR, incident volume, recurring failures, and automation coverage.
• Experience leading geographically distributed technical teams and implementing effective on-call, escalation, and coverage models.
• Capability to lead collaboration across Engineering, Product, IT, implementation, and customer-facing organizations during production incidents and initiatives focused on reliability.
• Preferred: Background in operating high-volume healthcare, financial services, SaaS, or other regulated production data environments.
• Preferred: Familiarity with MuleSoft, Flatfile, or similar enterprise integration and data ingestion platforms.
• Preferred: Experience with CI/CD tooling, cloud observability platforms, workflow orchestration, and automated recovery strategies.
• Preferred: Experience utilizing AI-assisted tools for production support, incident analysis, documentation, or operational automation.
• Preferred: Knowledge in healthcare payer data, CMS submissions, EDI/X12, enrollment, claims, risk adjustment, or related domains.
• Competitive salary.
• Medical, Dental, and Vision benefits.
• 401k match.
• Generous PTO plan.
Mercor
Mission Lane
ICF
The Cigna Group
Get handpicked remote jobs straight to your inbox weekly.