
Senior AI Infrastructure Engineer – Hardware Infrastructure Engineering
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in California, +4 more states.
• Develop and manage scalable telemetry pipelines for metrics, logs, traces, and events across on-premise, CSP, and NCP clusters.
• Create standardized instrumentation, collection, storage, and access protocols to ensure consistent telemetry generation and usage.
• Provide dashboards, alerting systems, and analytical tools that enhance service visibility, detection, and troubleshooting capabilities.
• Standardize and automate workflows for incidents, maintenance, service on-call, and support on-call across HWInf.
• Integrate operational data and lifecycle signals to enhance ownership, escalation, communication, and learning post-incident.
• Develop reporting and AI-assisted tools to minimize manual work and enhance operational responsiveness.
• Build and sustain physical hardware and software catalogs as reliable sources of truth for infrastructure inventory, service ownership, dependencies, and documentation.
• Design consistent data models and integration pipelines that connect clusters, hardware, services, teams, and operational workflows.
• Enable self-service discovery features so engineers can determine what they operate, who is responsible for it, and how to support it.
• Bachelor's degree in Computer Science, Computer Engineering, or a related technical discipline, or equivalent experience.
• Over 8 years of experience in infrastructure security, platform engineering, or security tooling.
• Proficiency in one or more programming languages, including Python, Go, Typescript, or Java.
• Strong grasp of software and infrastructure principles, with practical experience applying them in production settings.
• Capability to lead cross-functional initiatives involving internal teams and external partners across engineering, product, finance, and security.
• Experience in building and managing incident-management processes using both internally developed and externally provided SaaS tools.
• Background in constructing, deploying, and maintaining machine learning models within production systems.
• Familiarity with AI agent frameworks or orchestration tools.
• Experience in building and operating modern observability platforms for scalable metrics, logs, traces, and profiling.
• Experience with service catalog and configuration management databases (CMDB).
• Equity
• Benefits
LocalStack
Coinbase
Health Care Service Corporation
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.