Remotery

Senior AI Infrastructure Engineer – Platform Operations

Posted Aug 3

This is a fully remote position, open to applicants in Europe.

📋 Description

• Oversee the investigation and resolution of intricate incidents related to infrastructure, networking, and platforms.

• Serve as a senior escalation point for operational teams during critical service-impacting situations.

• Provide support for expansive NVIDIA GPU infrastructure and high-performance networking settings.

• Diagnose complex issues involving Linux, Kubernetes, networking, storage, and hardware.

• Evaluate platform performance, capacity, stability, and reliability trends to proactively pinpoint risks.

• Lead root cause analysis efforts and implement long-term corrective measures.

• Collaborate with engineering teams, hardware suppliers, and datacenter staff to tackle complex technical problems.

• Engage in major incident management and service restoration efforts.

• Offer technical guidance for Kubernetes platform operations and supporting infrastructure services.

• Propel enhancements in platform reliability, observability, monitoring, and operational workflows.

• Discover opportunities to automate repetitive operational tasks and boost operational efficiency.

• Contribute to operational readiness assessments, infrastructure modifications, upgrades, and service launches.

• Assist in the adoption and functioning of AI-driven infrastructure services and operational capabilities via k0rdent AI.

• Assess emerging technologies and operational practices to enhance service delivery and platform resilience.

• Mentor and support AI Infrastructure & Platform Operations Engineers.

• Disseminate technical expertise through documentation, training sessions, and operational reviews.

• Develop and uphold operational standards, runbooks, troubleshooting manuals, and best practices.

• Help shape operational processes, escalation routes, and service reliability benchmarks.

• Act as a trusted technical consultant during operational planning and service enhancement initiatives.


⛳️ Requirements

• Over 7 years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or similar technical positions.

• Advanced Linux administration and troubleshooting capabilities.

• Solid networking skills, including the ability to diagnose complex performance, connectivity, and reliability challenges.

• Extensive experience operating Kubernetes in production settings.

• Experience in supporting large-scale production infrastructure and distributed systems.

• Proven track record of leading technical investigations and managing complex incidents.

• Experience in conducting root cause analysis and implementing long-term operational enhancements.

• Strong grasp of observability, monitoring, and service reliability methodologies.

• Exceptional troubleshooting and analytical abilities across diverse infrastructure domains.

• Strong communication, collaboration, and stakeholder management capabilities.

• Experience in one or more of the following areas is highly desirable: NVIDIA GPU infrastructure and accelerated computing platforms.

• InfiniBand networking and NVIDIA UFM.

• AI infrastructure environments.

• HPC environments.

• Platform Engineering or Site Reliability Engineering (SRE).

• Large-scale Kubernetes operations.

• Infrastructure automation technologies and Infrastructure-as-Code methodologies.

• Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.

• Performance analysis and optimization of distributed infrastructure platforms.

• Technical leadership, mentoring, or team lead roles.


🏝️ Benefits

• Work with some of the most advanced AI infrastructure environments currently in production.

• Engage with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.

• Contribute to the establishment of operational standards and reliability practices for next-generation AI infrastructure services.

• Influence the adoption of AI-driven operational capabilities through k0rdent AI.

• Collaborate with highly skilled engineers in solving complex infrastructure and platform challenges at scale.

• Become part of a growing organization that is heavily investing in AI infrastructure, platform services, and operational innovation.

People also viewed

Quantiphi5 hours ago

Platform Engineer

CA flagCanada, +1 more countryFull-timePlatform Engineer
ApplyView job
Encompass Corporation5 hours ago

Senior Platform Engineer

RS flagSerbia OnlyFull-timePlatform Engineer
ApplyView job
Shippit7 hours ago

Senior Platform Engineer

ID flagIndonesia OnlyFull-timePlatform Engineer
ApplyView job
Group O7 hours ago

Associate Power Platform Developer

US flagIllinois OnlyFull-timePlatform Engineer$65k – $75k/year
ApplyView job
ServiceTitan10 hours ago

Staff Software Engineer, Platform – DevEx

US flagCalifornia, +8 more statesFull-timePlatform Engineer$185.3k – $297.4k/year
ApplyView job
New Era Technology10 hours ago

Senior QRadar Platform Engineer

US flagUnited States OnlyFull-timePlatform Engineer$80 – $125/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers