
AI Infrastructure & Platform Operations Engineer
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in California.
• Oversee, manage, and assist with production AI infrastructure platforms.
• Investigate and resolve incidents related to infrastructure, networking, hardware, and platforms.
• Provide support for NVIDIA GPU infrastructure and associated platform services.
• Monitor and troubleshoot environments based on Kubernetes.
• Analyze performance, availability, and reliability issues among infrastructure and platform components.
• Collaborate with engineering teams, hardware vendors, data center staff, and service delivery teams to address technical challenges.
• Engage in incident response, root cause analysis, and initiatives for operational enhancement.
• Contribute to advancements in monitoring, observability, automation, and operational processes.
• Maintain documentation, runbooks, and knowledge articles for operational procedures.
• A minimum of 3 years of experience in infrastructure operations, platform operations, network operations, site reliability engineering, cloud operations, datacenter operations, or similar technical positions.
• Proficient in Linux administration and troubleshooting.
• Solid understanding of networking concepts with experience in diagnosing infrastructure-related challenges.
• Familiarity with Kubernetes in production settings.
• Experience in supporting production infrastructure and services.
• Strong analytical and problem-solving capabilities.
• Background in structured operational and incident management processes.
• Exceptional communication and teamwork skills.
• Ability to operate within a shift-based work environment.
• Collaborate with a well-established leader in the cloud infrastructure sector based in Silicon Valley.
• Work alongside exceptionally passionate, skilled, and engaging colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies.
• Participate in cutting-edge, open-source innovation.
• Flourish in a dynamic environment of a young company that values openness, collaboration, risk-taking, and continuous growth.
• Opportunities for professional development and training.
• Attend conferences and working groups.
• Enjoy company outings, happy hours, hackathons, and tech talks.
• Receive a competitive compensation package along with a robust benefits plan.
Learning Technologies Group plc
Autodesk
NVIDIA
Astreya
Get handpicked remote jobs straight to your inbox weekly.