
Senior AI Infrastructure, Platform Operations Engineer
Posted Jul 21

Posted Jul 21
This is a fully remote position, open to applicants in United States.
• Lead the investigation and resolution of intricate incidents related to infrastructure, networking, and platforms.
• Serve as a senior escalation resource for operational teams during critical events that affect service.
• Support expansive NVIDIA GPU infrastructure and high-performance networking environments.
• Diagnose and resolve complex issues involving Linux, Kubernetes, networking, storage, and hardware.
• Analyze trends in platform performance, capacity, stability, and reliability to proactively identify potential risks.
• Spearhead root cause analysis initiatives and drive long-term corrective measures.
• Collaborate with engineering teams, hardware vendors, and datacenter personnel to tackle intricate technical challenges.
• Provide technical leadership for Kubernetes platform operations and associated infrastructure services.
• Promote enhancements in platform reliability, observability, monitoring, and operational procedures.
• Identify chances to automate repetitive operational tasks and enhance operational efficiency.
• Contribute to operational readiness assessments, infrastructure modifications, upgrades, and service rollouts.
• Support the implementation and operation of AI-driven infrastructure services and capabilities via k0rdent AI.
• Assess emerging technologies and operational methodologies to enhance service delivery and platform resilience.
• Mentor and assist AI Infrastructure & Platform Operations Engineers.
• Disseminate technical knowledge through documentation, training sessions, and operational evaluations.
• Develop and maintain operational standards, runbooks, troubleshooting guides, and best practices.
• Act as a trusted technical advisor during operational planning and initiatives aimed at service improvement.
• Over 7 years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or similar technical roles.
• Advanced Linux administration and troubleshooting skills.
• In-depth networking expertise, including the ability to diagnose complex performance, connectivity, and reliability challenges.
• Extensive experience operating Kubernetes in production settings.
• Background in supporting large-scale production infrastructure and distributed systems.
• Proven track record of leading technical investigations and managing complex incidents.
• Experience in performing root cause analysis and implementing long-term operational enhancements.
• Strong grasp of observability, monitoring, and service reliability practices.
• Excellent troubleshooting and analytical abilities across various infrastructure domains.
• Strong communication, collaboration, and stakeholder management skills.
• Work with a prominent Silicon Valley leader in the cloud infrastructure sector.
• Collaborate with exceptionally passionate, talented, and engaging colleagues who assist Fortune 500 and Global 2000 customers in implementing next-generation cloud technologies.
• Participate in cutting-edge, open-source innovation.
• Excel in a dynamic environment of a young company that values openness, collaboration, risk-taking, and continuous growth.
• Opportunities for professional development and training.
• Attend conferences and working groups.
• Enjoy company outings, happy hours, hackathons, and tech discussions.
• Receive a competitive compensation package along with a robust benefits program.
Quantiphi
Encompass Corporation
Shippit
Group O
Get handpicked remote jobs straight to your inbox weekly.