
AI Infrastructure & Platform Operations Engineer
Posted Jul 7

Posted Jul 7
This is a fully remote position, open to applicants in Europe.
β’ Oversee, manage, and provide support for production AI infrastructure platforms.
β’ Analyze and resolve incidents related to infrastructure, networking, hardware, and platforms.
β’ Provide support for NVIDIA GPU infrastructure and associated platform services.
β’ Monitor and diagnose issues in Kubernetes-based environments.
β’ Examine performance, availability, and reliability challenges across infrastructure and platform elements.
β’ Work collaboratively with engineering teams, hardware vendors, datacenter staff, and service delivery teams to address technical challenges.
β’ Engage in incident response, root cause analysis, and initiatives for operational enhancement.
β’ Contribute to advancements in monitoring, observability, automation, and operational processes.
β’ Keep operational documentation, runbooks, and knowledge articles updated.
β’ Minimum of 3 years' experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or similar technical roles.
β’ Proficient in Linux administration and troubleshooting.
β’ Solid understanding of networking principles and experience in diagnosing infrastructure-related problems.
β’ Familiarity with Kubernetes in production settings.
β’ Experience in supporting production infrastructure and services.
β’ Strong analytical and problem-solving capabilities.
β’ Background in structured operational and incident management processes.
β’ Exceptional communication and collaboration abilities.
β’ Willingness to work in a shift-based operational setting.
β’ Experience in one or more of the following areas is highly advantageous: NVIDIA GPU infrastructure and accelerated computing platforms.
β’ InfiniBand networking and NVIDIA UFM.
β’ Operations on Kubernetes platforms.
β’ AI infrastructure or HPC environments.
β’ Site Reliability Engineering (SRE) or Platform Engineering.
β’ Familiarity with observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.
β’ Knowledge of infrastructure automation technologies and Infrastructure-as-Code practices.
β’ Experience with large-scale distributed systems and production platforms.
β’ Work with some of the most advanced AI infrastructure environments currently in production.
β’ Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking settings.
β’ Contribute to defining the operational and support framework for next-generation AI infrastructure.
β’ Join a team that is shaping the future of AI-powered operations through k0rdent AI.
β’ Become part of a growing organization that is making significant investments in AI infrastructure and platform services.
PSI CRO AG
PSI CRO AG
PSI CRO AG
PSI CRO AG
Get handpicked remote jobs straight to your inbox weekly.