
Senior AI Infrastructure, Platform Operations Engineer
Posted Jul 7

Posted Jul 7
This is a fully remote position, open to applicants in Europe.
• Oversee the investigation and resolution of intricate incidents related to infrastructure, networking, and platforms.
• Serve as a senior point of escalation for operational teams during significant service-impacting situations.
• Provide support for extensive NVIDIA GPU infrastructure and high-performance networking settings.
• Diagnose complex issues involving Linux, Kubernetes, networking, storage, and hardware.
• Examine platform performance, capacity, stability, and reliability trends to proactively pinpoint risks.
• Lead root cause analysis initiatives and implement long-term corrective measures.
• Collaborate with engineering teams, hardware vendors, and datacenter staff to tackle intricate technical challenges.
• Engage in major incident management and service restoration efforts.
• Offer technical leadership for Kubernetes platform operations and associated infrastructure services.
• Promote enhancements in platform reliability, observability, monitoring, and operational procedures.
• Identify chances to automate repetitive operational tasks and enhance operational efficiency.
• Contribute to operational readiness assessments, infrastructure modifications, upgrades, and service rollouts.
• Support the integration and functioning of AI-powered infrastructure services and operational capabilities using k0rdent AI.
• Assess emerging technologies and operational methodologies to enhance service delivery and platform resilience.
• Mentor and assist AI Infrastructure & Platform Operations Engineers.
• Disseminate technical knowledge through documentation, training sessions, and operational reviews.
• Create and maintain operational standards, runbooks, troubleshooting guides, and best practices.
• Assist in defining operational processes, escalation routes, and service reliability benchmarks.
• Act as a trusted technical advisor in operational planning and service enhancement initiatives.
• A minimum of 7 years’ experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or related technical roles.
• Expertise in Linux administration and troubleshooting.
• Strong networking skills, including the ability to diagnose complex performance, connectivity, and reliability issues.
• Significant experience managing Kubernetes in production environments.
• Experience supporting large-scale production infrastructure and distributed systems.
• Proven track record in leading technical investigations and managing complex incidents.
• Experience in conducting root cause analyses and promoting long-term operational enhancements.
• Robust understanding of observability, monitoring, and service reliability strategies.
• Exceptional troubleshooting and analytical abilities across various infrastructure domains.
• Strong communication, collaboration, and stakeholder management skills.
• Experience in one or more of the following areas is highly desirable: NVIDIA GPU infrastructure and accelerated computing platforms, InfiniBand networking and NVIDIA UFM, AI infrastructure environments, HPC environments, Platform Engineering or Site Reliability Engineering (SRE), large-scale Kubernetes operations, infrastructure automation technologies and Infrastructure-as-Code practices, observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry, performance analysis and optimization of distributed infrastructure platforms, and technical leadership, mentoring, or team lead responsibilities.
• Work within some of the most advanced AI infrastructure environments currently in production.
• Collaborate with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
• Help establish operational standards and reliability practices for next-generation AI infrastructure services.
• Influence the integration of AI-powered operational capabilities through k0rdent AI.
• Work alongside highly skilled engineers addressing complex infrastructure and platform challenges at scale.
• Join a growing organization that is heavily investing in AI infrastructure, platform services, and operational innovation.
PSI CRO AG
PSI CRO AG
PSI CRO AG
PSI CRO AG
Get handpicked remote jobs straight to your inbox weekly.