
Senior AI Infrastructure Engineer – Platform Operations
Posted Aug 3

Posted Aug 3
This is a fully remote position, open to applicants in Europe.
• Oversee the investigation and resolution of intricate incidents related to infrastructure, networking, and platforms.
• Serve as a senior escalation point for operational teams during critical service-impacting situations.
• Provide support for expansive NVIDIA GPU infrastructure and high-performance networking settings.
• Diagnose complex issues involving Linux, Kubernetes, networking, storage, and hardware.
• Evaluate platform performance, capacity, stability, and reliability trends to proactively pinpoint risks.
• Lead root cause analysis efforts and implement long-term corrective measures.
• Collaborate with engineering teams, hardware suppliers, and datacenter staff to tackle complex technical problems.
• Engage in major incident management and service restoration efforts.
• Offer technical guidance for Kubernetes platform operations and supporting infrastructure services.
• Propel enhancements in platform reliability, observability, monitoring, and operational workflows.
• Discover opportunities to automate repetitive operational tasks and boost operational efficiency.
• Contribute to operational readiness assessments, infrastructure modifications, upgrades, and service launches.
• Assist in the adoption and functioning of AI-driven infrastructure services and operational capabilities via k0rdent AI.
• Assess emerging technologies and operational practices to enhance service delivery and platform resilience.
• Mentor and support AI Infrastructure & Platform Operations Engineers.
• Disseminate technical expertise through documentation, training sessions, and operational reviews.
• Develop and uphold operational standards, runbooks, troubleshooting manuals, and best practices.
• Help shape operational processes, escalation routes, and service reliability benchmarks.
• Act as a trusted technical consultant during operational planning and service enhancement initiatives.
• Over 7 years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or similar technical positions.
• Advanced Linux administration and troubleshooting capabilities.
• Solid networking skills, including the ability to diagnose complex performance, connectivity, and reliability challenges.
• Extensive experience operating Kubernetes in production settings.
• Experience in supporting large-scale production infrastructure and distributed systems.
• Proven track record of leading technical investigations and managing complex incidents.
• Experience in conducting root cause analysis and implementing long-term operational enhancements.
• Strong grasp of observability, monitoring, and service reliability methodologies.
• Exceptional troubleshooting and analytical abilities across diverse infrastructure domains.
• Strong communication, collaboration, and stakeholder management capabilities.
• Experience in one or more of the following areas is highly desirable: NVIDIA GPU infrastructure and accelerated computing platforms.
• InfiniBand networking and NVIDIA UFM.
• AI infrastructure environments.
• HPC environments.
• Platform Engineering or Site Reliability Engineering (SRE).
• Large-scale Kubernetes operations.
• Infrastructure automation technologies and Infrastructure-as-Code methodologies.
• Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
• Performance analysis and optimization of distributed infrastructure platforms.
• Technical leadership, mentoring, or team lead roles.
• Work with some of the most advanced AI infrastructure environments currently in production.
• Engage with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
• Contribute to the establishment of operational standards and reliability practices for next-generation AI infrastructure services.
• Influence the adoption of AI-driven operational capabilities through k0rdent AI.
• Collaborate with highly skilled engineers in solving complex infrastructure and platform challenges at scale.
• Become part of a growing organization that is heavily investing in AI infrastructure, platform services, and operational innovation.
Quantiphi
Encompass Corporation
Shippit
Group O
Get handpicked remote jobs straight to your inbox weekly.