
AI Infrastructure & Platform Operations Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Europe.
• Oversee, manage, and provide support for production AI infrastructure platforms.
• Identify and resolve incidents related to infrastructure, networking, hardware, and platform operations.
• Provide support for NVIDIA GPU infrastructure and its related platform services.
• Monitor and troubleshoot environments based on Kubernetes.
• Investigate issues concerning performance, availability, and reliability across infrastructure and platform components.
• Collaborate with engineering teams, hardware suppliers, datacenter staff, and service delivery teams to address technical challenges.
• Engage in incident response, root cause analysis, and initiatives for operational enhancement.
• Contribute to advancements in monitoring, observability, automation, and operational processes.
• Maintain operational documentation, runbooks, and knowledge articles.
• Minimum of 3 years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or similar technical roles.
• Proficient in Linux administration and troubleshooting.
• Solid understanding of networking concepts and experience in diagnosing infrastructure-related problems.
• Familiarity with Kubernetes in production settings.
• Experience in supporting production infrastructure and services.
• Strong analytical and problem-solving capabilities.
• Background in structured operational and incident management processes.
• Excellent communication and teamwork abilities.
• Capability to operate within a shift-based work environment.
• Experience in one or more of the following areas is highly desirable: NVIDIA GPU infrastructure and accelerated computing platforms.
• Knowledge of InfiniBand networking and NVIDIA UFM.
• Expertise in Kubernetes platform operations.
• Familiarity with AI infrastructure or HPC environments.
• Background in Site Reliability Engineering (SRE) or Platform Engineering.
• Experience with observability platforms like Grafana, Prometheus, ELK, or OpenTelemetry.
• Knowledge of infrastructure automation technologies and Infrastructure-as-Code practices.
• Experience with large-scale distributed systems and production platforms.
• Work with some of the most advanced AI infrastructure environments currently in production.
• Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
• Contribute to defining the operation and support of next-generation AI infrastructure.
• Be part of a team that is shaping the future of AI-powered operations through k0rdent AI.
• Join a growing organization that is making significant investments in AI infrastructure and platform services.
Blue Ocean Global Technology
NVIDIA
UFS Tech
ShiftKey
Get handpicked remote jobs straight to your inbox weekly.