Remotery

Senior AI Infrastructure, Platform Operations Engineer

Posted Jul 21

This is a fully remote position, open to applicants in United States.

📋 Description

• Lead the investigation and resolution of intricate incidents related to infrastructure, networking, and platforms.

• Serve as a senior escalation resource for operational teams during critical events that affect service.

• Support expansive NVIDIA GPU infrastructure and high-performance networking environments.

• Diagnose and resolve complex issues involving Linux, Kubernetes, networking, storage, and hardware.

• Analyze trends in platform performance, capacity, stability, and reliability to proactively identify potential risks.

• Spearhead root cause analysis initiatives and drive long-term corrective measures.

• Collaborate with engineering teams, hardware vendors, and datacenter personnel to tackle intricate technical challenges.

• Provide technical leadership for Kubernetes platform operations and associated infrastructure services.

• Promote enhancements in platform reliability, observability, monitoring, and operational procedures.

• Identify chances to automate repetitive operational tasks and enhance operational efficiency.

• Contribute to operational readiness assessments, infrastructure modifications, upgrades, and service rollouts.

• Support the implementation and operation of AI-driven infrastructure services and capabilities via k0rdent AI.

• Assess emerging technologies and operational methodologies to enhance service delivery and platform resilience.

• Mentor and assist AI Infrastructure & Platform Operations Engineers.

• Disseminate technical knowledge through documentation, training sessions, and operational evaluations.

• Develop and maintain operational standards, runbooks, troubleshooting guides, and best practices.

• Act as a trusted technical advisor during operational planning and initiatives aimed at service improvement.


⛳️ Requirements

• Over 7 years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or similar technical roles.

• Advanced Linux administration and troubleshooting skills.

• In-depth networking expertise, including the ability to diagnose complex performance, connectivity, and reliability challenges.

• Extensive experience operating Kubernetes in production settings.

• Background in supporting large-scale production infrastructure and distributed systems.

• Proven track record of leading technical investigations and managing complex incidents.

• Experience in performing root cause analysis and implementing long-term operational enhancements.

• Strong grasp of observability, monitoring, and service reliability practices.

• Excellent troubleshooting and analytical abilities across various infrastructure domains.

• Strong communication, collaboration, and stakeholder management skills.


🏝️ Benefits

• Work with a prominent Silicon Valley leader in the cloud infrastructure sector.

• Collaborate with exceptionally passionate, talented, and engaging colleagues who assist Fortune 500 and Global 2000 customers in implementing next-generation cloud technologies.

• Participate in cutting-edge, open-source innovation.

• Excel in a dynamic environment of a young company that values openness, collaboration, risk-taking, and continuous growth.

• Opportunities for professional development and training.

• Attend conferences and working groups.

• Enjoy company outings, happy hours, hackathons, and tech discussions.

• Receive a competitive compensation package along with a robust benefits program.

People also viewed

Quantiphi16 hours ago

Platform Engineer

CA flagCanada, +1 more countryFull-timePlatform Engineer
ApplyView job
Encompass Corporation16 hours ago

Senior Platform Engineer

RS flagSerbia OnlyFull-timePlatform Engineer
ApplyView job
Shippit18 hours ago

Senior Platform Engineer

ID flagIndonesia OnlyFull-timePlatform Engineer
ApplyView job
Group O18 hours ago

Associate Power Platform Developer

US flagIllinois OnlyFull-timePlatform Engineer$65k – $75k/year
ApplyView job
ServiceTitan21 hours ago

Staff Software Engineer, Platform – DevEx

US flagCalifornia, +8 more statesFull-timePlatform Engineer$185.3k – $297.4k/year
ApplyView job
New Era Technology21 hours ago

Senior QRadar Platform Engineer

US flagUnited States OnlyFull-timePlatform Engineer$80 – $125/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers