
ML Infrastructure Operations Engineer
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in Mexico.
• Set up and manage ML environments utilizing Kubernetes, Docker, and YAML configurations.
• Execute, monitor, and validate distributed ML training jobs across multi-node accelerator clusters.
• Review, modify, and run Python and Bash scripts to automate job execution and adjust runtime settings.
• Diagnose workload failures, gather logs, pinpoint infrastructure or configuration issues, and document results.
• Monitor workload execution status, update on testing progress, and work collaboratively with engineering teams to enhance reliability and operational efficiency.
• Over 3 years of experience in ML Operations, Software Testing, Systems QA, Linux Systems Administration, or a similar technical role.
• Strong command of Linux command-line operations and shell environments.
• Experience in reading and modifying Python and Bash scripts.
• Familiarity with Kubernetes, Docker, and containerized environments.
• Experience in executing and monitoring distributed workloads.
• Proficiency in English: C1
• Base salary
• Major Medical Expenses Insurance (includes dental and vision plan)
• 15 days of holiday bonus
• 25% vacation premium
• 12 days of vacation (starting from the first year)
• Social security
• Biweekly grocery vouchers
Hanwha Energy USA
Accenture Federal Services
PSI CRO AG
PSI CRO AG
Get handpicked remote jobs straight to your inbox weekly.