Remotery

AI Infrastructure & Platform Operations Engineer

atMirantisRemoteEuropeFull-timeInfrastructure EngineerMid-levelSenior$60k – $67k/year

Posted 1 day ago

This is a fully remote position, open to applicants in Europe.

📋 Description

• Oversee, manage, and provide support for production AI infrastructure platforms.

• Identify and resolve incidents related to infrastructure, networking, hardware, and platform operations.

• Provide support for NVIDIA GPU infrastructure and its related platform services.

• Monitor and troubleshoot environments based on Kubernetes.

• Investigate issues concerning performance, availability, and reliability across infrastructure and platform components.

• Collaborate with engineering teams, hardware suppliers, datacenter staff, and service delivery teams to address technical challenges.

• Engage in incident response, root cause analysis, and initiatives for operational enhancement.

• Contribute to advancements in monitoring, observability, automation, and operational processes.

• Maintain operational documentation, runbooks, and knowledge articles.


⛳️ Requirements

• Minimum of 3 years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or similar technical roles.

• Proficient in Linux administration and troubleshooting.

• Solid understanding of networking concepts and experience in diagnosing infrastructure-related problems.

• Familiarity with Kubernetes in production settings.

• Experience in supporting production infrastructure and services.

• Strong analytical and problem-solving capabilities.

• Background in structured operational and incident management processes.

• Excellent communication and teamwork abilities.

• Capability to operate within a shift-based work environment.

• Experience in one or more of the following areas is highly desirable: NVIDIA GPU infrastructure and accelerated computing platforms.

• Knowledge of InfiniBand networking and NVIDIA UFM.

• Expertise in Kubernetes platform operations.

• Familiarity with AI infrastructure or HPC environments.

• Background in Site Reliability Engineering (SRE) or Platform Engineering.

• Experience with observability platforms like Grafana, Prometheus, ELK, or OpenTelemetry.

• Knowledge of infrastructure automation technologies and Infrastructure-as-Code practices.

• Experience with large-scale distributed systems and production platforms.


🏝️ Benefits

• Work with some of the most advanced AI infrastructure environments currently in production.

• Gain exposure to NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.

• Contribute to defining the operation and support of next-generation AI infrastructure.

• Be part of a team that is shaping the future of AI-powered operations through k0rdent AI.

• Join a growing organization that is making significant investments in AI infrastructure and platform services.

People also viewed

Blue Ocean Global Technology1 day ago

Senior Staff Full Stack Software Engineer – Infrastructure Platforms

HU flagHungary OnlyFull-timeInfrastructure Engineer€61.5k – €76.9k/year
ApplyView job
NVIDIA1 day ago

Senior Platform Engineer, Network Infrastructure

IN flagIndia OnlyFull-timeInfrastructure Engineer
ApplyView job
UFS Tech1 day ago

Lead Platform Infrastructure Engineer

US flagUnited States OnlyFull-timeInfrastructure Engineer
ApplyView job
ShiftKey1 day ago

Staff AI Engineer – AI Infrastructure, Agentic Platform

PL flagPoland OnlyFull-timeInfrastructure Engineer
ApplyView job
nbn® Australia1 day ago

Senior Field Specialist – Critical Infrastructure

AU flagAustralia OnlyFull-timeInfrastructure Engineer
ApplyView job
CoLogix Analytics1 day ago

AI Infrastructure Engineer – Emerging Technologies

US flagOhio, +2 more statesFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers