
Director of Platform Engineering
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in United States.
• Evaluate and consolidate CI/CD and deployment methodologies across various product lines and manage their implementation.
• Actively participate in containerization and the automation of deployment processes.
• Enhance deployment frequency, predictability, consistency across environments, and the safety of releases.
• Establish Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and standards for production reliability.
• Manage standardized monitoring, alerting, observability patterns, and dashboards.
• Collaborate with development teams to ensure consistent instrumentation.
• Oversee company-wide on-call tools, design rotation schedules, and create escalation policies.
• Define and enforce Root Cause Analysis (RCA) and post-incident review practices, utilizing AI to expedite investigations.
• Cultivate widely accepted operational practices across teams through influence without direct authority.
• Ensure the reliability of core production databases, focusing on uptime, query performance, and capacity management.
• Establish governance for schema changes and review protocols for a high-traffic, monolithic database.
• Collaborate with Infrastructure on backup protocols, continuity plans, and Business Continuity and Disaster Recovery (BCDR) testing.
• Define operational metrics and develop dashboards and reports for engineering and executive leadership.
• Formalize an AI-driven natural-language query interface for observability data.
• Partner on shared component libraries, frontend build tools, and contracts related to design systems/APIs.
• Act as a sponsor for a dotted-line Frontend Guild.
• Identify structural patterns and architectural shifts across products and drive associated enhancements.
• Influence cross-functional platform architecture without having ownership of product or feature architecture.
• Construct the structure and hiring strategy for the function from the ground up.
• Establish the operating cadence for the team, including on-call duties, incident reviews, and roadmap planning.
• Report directly to the SVP of Engineering and represent platform reliability and operational maturity to the Executive Leadership Team (ELT).
• Proven experience in modernizing CI/CD and deployment pipelines at scale.
• Formal experience in Site Reliability Engineering (SRE), including SLOs, observability/alerting tools, and incident management.
• Experience in designing on-call and incident-management programs, such as PagerDuty or similar, and enforcing RCA practices.
• Knowledge in database reliability and operations, including schema governance, query optimization, backup, and BCDR for a high-traffic production database, or direct management of an individual with such experience.
• Experience in vertical SaaS, regulated environments, or mission-critical software is a plus.
• Background in public safety is advantageous.
• Executive presence to discuss matters of reliability, risk, and operational maturity with the ELT in business-oriented terms.
• A hands-on approach to working in pipelines, dashboards, or incident channels.
• Authorization to work for any employer in the United States.
• No visa sponsorship or transition support for employment is available.
• Successful completion of a criminal background check is required.
• Must complete I-9 employment authorization verification through E-Verify.
• Medical, dental, and vision insurance.
• Flexible Spending Account (FSA) and Health Savings Account (HSA).
• 401(k) retirement plan.
• Flexible Paid Time Off (PTO).
• Fully remote work environment.
• Technology stipend.
• Opportunities for career advancement.
• Competitive salary.
• Minimal travel requirements.
• Reasonable accommodations during employment and the interview process.
Agility Technologies Inc
American College of Education
Faire
Faire
Get handpicked remote jobs straight to your inbox weekly.