
Senior Director, Platform & AI Infrastructure
Posted Jul 20

Posted Jul 20
This is a fully remote position, open to applicants in United States.
• Take ownership of the AI/ML platform, including GPU capacity strategy, model serving and inference latency, training and fine-tuning infrastructure, MLOps and evaluation pipelines, vector and feature stores, as well as the RAG and agentic patterns utilized by our product teams.
• Develop the incident response program, which encompasses on-call structure, severity definitions, incident command, communication protocols, postmortems, and ensuring systemic fixes are implemented.
• Enhance the observability stack across metrics, logs, traces, and synthetics, while establishing SLOs and reporting on them.
• Manage the Azure and OCI infrastructure footprint, focusing on infrastructure-as-code, capacity planning, and ensuring reliability for both CPU and GPU workloads.
• Oversee operations, performance, high availability/disaster recovery, and roadmap for databases such as Oracle, SQL Server, MongoDB, and similar technologies.
• Provide briefings to executives regarding reliability and risk; lead internal communications during incidents; and engage with customers when significant incidents impact them.
• Lead a globally distributed team comprising managers and senior individual contributors, fostering a strong cultural environment and leadership pipeline.
• Over 10 years of experience in platform engineering, site reliability engineering, infrastructure, or AI/ML infrastructure, with at least 5 years in leadership roles.
• Proven production experience managing ML/AI workloads at scale, particularly with GPU infrastructure, model serving, MLOps, or LLM/inference platforms.
• Knowledge of the contemporary AI stack, including vector databases, RAG, agent frameworks, evaluation, and the considerations of build versus buy among foundation model providers and open-source solutions.
• Experience in building or revamping an incident response or observability program at a large scale.
• Demonstrated measurable improvements in reliability metrics (MTTR, availability, change failure rate) within a cloud environment.
• Strong communication skills, capable of engaging with executives, the Board, customers, and the organization during incidents.
• Familiarity with zero-trust and modern identity platforms.
• Medical, dental, vision, life, and long-term disability insurance
• Health Savings Account (HSA)
• 401(k) retirement plan
Arctiq
Cisco
Prove
Hello Heart
Get handpicked remote jobs straight to your inbox weekly.