
Director of Platform Operations
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in United Kingdom.
• Take ownership of Arbor’s operational backbone across its suite of applications.
• Integrate Site Reliability Engineering, Security Engineering, and Developer Experience under a unified leadership structure.
• Ensure commitment to availability and accountability for the 99.9% availability SLA.
• Define, publish, and manage SLO and error-budget frameworks.
• Lead initiatives in observability, capacity planning, performance engineering, resilience testing, and reduction of toil.
• Communicate availability, latency, and error-budget consumption metrics to product teams, R&D leadership, executives, and key stakeholders.
• Drive improvements in systemic reliability and address recurring issues, isolated points of failure, and architectural vulnerabilities.
• Establish and manage vulnerability management for application code, dependencies, containers, and infrastructure.
• Develop vulnerability reporting and dashboards while driving remediation efforts against SLAs.
• Integrate security measures into engineering workflows via automated scanning, secure-by-default platform patterns, dependency hygiene, and team enablement.
• Collaborate with security leadership and the CISO on security posture, roadmap, compliance responsibilities, and audit documentation.
• Guide the transition to a “build it, run it” model, granting product engineering teams ownership of production, including on-call duties, alerting, and operational health.
• Oversee the delivery of the internal developer platform, golden paths, self-service tools, CI/CD processes, and environment provisioning.
• Manage the platform as a product, engaging with internal customers, establishing service levels, tracking adoption metrics, gathering feedback, and maintaining an impact-focused roadmap.
• Oversee engineering productivity metrics such as deployment frequency, lead time for change, change failure rate, and time to restore.
• Proven experience in senior leadership roles within SRE, platform, or infrastructure functions at scale.
• Proven track record of maintaining availability against a defined SLA in a customer-facing SaaS setting.
• Extensive hands-on knowledge of SLIs, SLOs, error budgets, observability, capacity planning, and problem management.
• Experience managing major incident processes, including incident command frameworks, on-call design, and blameless post-incident reviews.
• Demonstrated success in materially improving MTTR.
• Experience in managing or closely collaborating on security posture, vulnerability management at scale, remediation SLAs, and executive or board-level reporting.
• Proven ability to lead a shift towards distributed production ownership (“build it, run it”).
• Familiarity with internal developer platforms or platform-as-a-product models.
• Experience leveraging engineering productivity metrics to inform investment decisions.
• Judgement in applying AI within operational workflows.
• Strong track record in building and developing engineering leaders, including effective performance management.
• Capacity to achieve results through influence across teams outside of direct reporting lines.
• Strong technical credibility with AWS and Terraform.
• Exceptional communication and influencing skills, including the ability to present technical and risk concepts to non-technical audiences during live incidents.
• Desirable: experience in enterprise SaaS at scale, preferably in EdTech or other data-sensitive or regulated sectors.
• Desirable: knowledge of ISO 27001, SOC 2, Cyber Essentials Plus, and UK GDPR compliance obligations.
• Desirable: hands-on experience with AIOps or AI-assisted incident management tools.
• Desirable: experience managing multi-product, multi-market environments based on shared platform foundations.
• Desirable: familiarity with PHP-based estates, Docker, containerization, Kanban, and agile methodologies.
• Desirable: accountability for FinOps and cloud cost efficiency.
• Must be eligible to work without visa sponsorship, as sponsorship is not available.
• A dedicated wellbeing team and initiatives including mindfulness, lunch and learns, manager training, and mental health first aid training.
• 32 days of holiday plus Bank Holidays (25 days annual leave plus 7 company-wide days).
• Life Assurance valued at 3x annual salary.
• Comprehensive wellness benefits through AIG Smart Health, offering 24/7 virtual GP access, mental health support, counselling, and personalized health checks.
• Private Dental Insurance provided by Bupa.
• Salary sacrifice pension scheme offered by Scottish Widows.
• Enhanced maternity and adoption leave (20 weeks full pay).
• Enhanced paternity leave (6 weeks full pay).
• 5 complimentary return-to-work maternity coaching sessions.
• Access to Calm and Bippit for financial wellbeing coaching.
• Flexible working arrangements.
• Social committees and events at team, office, and company-wide levels.
• A dedicated professional development training budget, including CPD courses, upskilling resources, and professional memberships.
• One paid volunteer day each year.
• Dog-friendly office environment.
• Referral voucher worth up to £200.
CareMetx, LLC
Zocdoc
Get handpicked remote jobs straight to your inbox weekly.