
AI Infrastructure Engineer
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in Japan.
• Oversee and manage production AI infrastructure environments.
• Lead incident response and troubleshooting initiatives, ensuring timely restoration of services during outages or performance issues.
• Diagnose infrastructure and networking problems across bare-metal and/or cloud environments involving multiple vendors.
• Perform root cause analysis and promote product and operational enhancements.
• Contribute to and enhance operational documentation and the knowledge base.
• Collaborate with global team members across various time zones to maintain continuous operational coverage, including occasional work on weekends and holidays.
• Demonstrated experience in managing and operating large-scale production systems (bare-metal and/or cloud).
• Strong working knowledge of Kubernetes with excellent and demonstrable troubleshooting abilities.
• Experience in configuring, customizing, and extending logging and monitoring tools (e.g., Prometheus, Grafana, ELK, or similar).
• Familiarity with infrastructure automation technologies and Infrastructure-as-Code practices (e.g., Ansible, Terraform, or similar).
• Proficient verbal and written communication skills in English.
• Strong analytical and problem-solving capabilities, adept at navigating complex and ambiguous technical challenges.
• Willingness to occasionally work during weekends and holidays.
• Nice to have: Prior experience in building, scaling, and managing High-Performance Computing (HPC) environments.
• Hands-on experience managing large-scale Kubernetes platforms in production.
• A solid understanding of NVIDIA GPU technologies and their associated software stack.
• Proficiency in scripting languages (e.g., Python, Bash, Go).
• Collaborate with a leading Silicon Valley company in the cloud infrastructure sector;
• Work alongside exceptionally passionate, talented, and engaging colleagues, assisting Fortune 500 and Global 2000 clients in implementing next-generation cloud technologies;
• Participate in cutting-edge, open-source innovation;
• Flourish in a high-energy environment of a dynamic company that values openness, collaboration, risk-taking, and ongoing growth;
• Access to professional development and training opportunities;
• Attend conferences and working groups;
• Enjoy company outings, happy hours, hackathons, and tech talks;
• Receive a competitive compensation package paired with a robust benefits plan.
Teleperformance
Trilon Group
Carbon60
fal
Get handpicked remote jobs straight to your inbox weekly.