Remotery

Senior Director, Platform & AI Infrastructure

Posted Jul 20

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the AI/ML platform, including GPU capacity strategy, model serving and inference latency, training and fine-tuning infrastructure, MLOps and evaluation pipelines, vector and feature stores, as well as the RAG and agentic patterns utilized by our product teams.

• Develop the incident response program, which encompasses on-call structure, severity definitions, incident command, communication protocols, postmortems, and ensuring systemic fixes are implemented.

• Enhance the observability stack across metrics, logs, traces, and synthetics, while establishing SLOs and reporting on them.

• Manage the Azure and OCI infrastructure footprint, focusing on infrastructure-as-code, capacity planning, and ensuring reliability for both CPU and GPU workloads.

• Oversee operations, performance, high availability/disaster recovery, and roadmap for databases such as Oracle, SQL Server, MongoDB, and similar technologies.

• Provide briefings to executives regarding reliability and risk; lead internal communications during incidents; and engage with customers when significant incidents impact them.

• Lead a globally distributed team comprising managers and senior individual contributors, fostering a strong cultural environment and leadership pipeline.


⛳️ Requirements

• Over 10 years of experience in platform engineering, site reliability engineering, infrastructure, or AI/ML infrastructure, with at least 5 years in leadership roles.

• Proven production experience managing ML/AI workloads at scale, particularly with GPU infrastructure, model serving, MLOps, or LLM/inference platforms.

• Knowledge of the contemporary AI stack, including vector databases, RAG, agent frameworks, evaluation, and the considerations of build versus buy among foundation model providers and open-source solutions.

• Experience in building or revamping an incident response or observability program at a large scale.

• Demonstrated measurable improvements in reliability metrics (MTTR, availability, change failure rate) within a cloud environment.

• Strong communication skills, capable of engaging with executives, the Board, customers, and the organization during incidents.

• Familiarity with zero-trust and modern identity platforms.


🏝️ Benefits

• Medical, dental, vision, life, and long-term disability insurance

• Health Savings Account (HSA)

• 401(k) retirement plan

People also viewed

ArctiqJul 26

ITSM Platform Developer – Halo

US flagPennsylvania OnlyFreelancePlatform Engineer$60 – $65/hour
ApplyView job
CiscoJul 26

Principal Engineer – Platform Engineering

US flagCalifornia, +3 more statesFull-timePlatform Engineer$231.4k – $331.8k/year
ApplyView job
ProveJul 25

Senior Software Engineer, Identity Platform

US flagUnited States OnlyFull-timePlatform Engineer$190k – $210k/year
ApplyView job
Hello HeartJul 25

Senior Director, Device Platform Engineering

US flagUnited States OnlyFull-timePlatform Engineer$220k – $238k/year
ApplyView job
AutodeskJul 25

Principal Machine Learning Developer – AI/ML Platform

CA flagCanada OnlyFull-timePlatform Engineer$153k – $224.4k/year
ApplyView job
ShiftKeyJul 25

Director – Platform & Infrastructure

US flagUnited States OnlyFull-timePlatform Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers