
AI Infrastructure & Platform Operations Engineer
Posted Jul 7

Posted Jul 7
This is a fully remote position, open to applicants in Europe.
β’ Oversee, manage, and provide support for production AI infrastructure platforms.
β’ Identify and resolve incidents related to infrastructure, networking, hardware, and platform operations.
β’ Provide support for NVIDIA GPU infrastructure and its related platform services.
β’ Monitor and troubleshoot environments based on Kubernetes.
β’ Investigate issues concerning performance, availability, and reliability across infrastructure and platform components.
β’ Collaborate with engineering teams, hardware suppliers, datacenter staff, and service delivery teams to address technical challenges.
β’ Engage in incident response, root cause analysis, and initiatives for operational enhancement.
β’ Contribute to advancements in monitoring, observability, automation, and operational processes.
β’ Maintain operational documentation, runbooks, and knowledge articles.
β’ Minimum of 3 years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or similar technical roles.
β’ Proficient in Linux administration and troubleshooting.
β’ Solid understanding of networking concepts and experience in diagnosing infrastructure-related problems.
β’ Familiarity with Kubernetes in production settings.
β’ Experience in supporting production infrastructure and services.
β’ Strong analytical and problem-solving capabilities.
β’ Background in structured operational and incident management processes.
β’ Excellent communication and teamwork abilities.
β’ Capability to operate within a shift-based work environment.
β’ Experience in one or more of the following areas is highly desirable: NVIDIA GPU infrastructure and accelerated computing platforms.
β’ Knowledge of InfiniBand networking and NVIDIA UFM.
β’ Expertise in Kubernetes platform operations.
β’ Familiarity with AI infrastructure or HPC environments.
β’ Background in Site Reliability Engineering (SRE) or Platform Engineering.
β’ Experience with observability platforms like Grafana, Prometheus, ELK, or OpenTelemetry.
β’ Knowledge of infrastructure automation technologies and Infrastructure-as-Code practices.
β’ Experience with large-scale distributed systems and production platforms.
β’ Work with some of the most advanced AI infrastructure environments currently in production.
β’ Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
β’ Contribute to defining the operation and support of next-generation AI infrastructure.
β’ Be part of a team that is shaping the future of AI-powered operations through k0rdent AI.
β’ Join a growing organization that is making significant investments in AI infrastructure and platform services.
PSI CRO AG
PSI CRO AG
PSI CRO AG
PSI CRO AG
Get handpicked remote jobs straight to your inbox weekly.