
Staff ML Engineer – AWS Trainium, SageMaker
Posted Sep 16

Posted Sep 16
This is a fully remote position, open to applicants in United States.
• Train and manage models using Amazon SageMaker, utilizing AWS Trainium as the foundational computing resource.
• Develop and enhance PyTorch training scripts, demonstrating an understanding of NeuronCore architecture, compiler behavior, memory management, and throughput trade-offs.
• Identify and troubleshoot issues during training runs that are attributed to hardware, effectively differentiating between data or code-related problems and compiler or device-level challenges.
• Convert Trainium job requests into efficient, cost-effective end-to-end training pipelines.
• Optimize distributed training sessions for both throughput and cost efficiency on SageMaker infrastructure.
• Collaborate directly with clients and internal engineering teams to define and execute production training workloads.
• Extensive hands-on experience with PyTorch, preferably including distributed or multi-device training capabilities.
• Practical experience with Amazon SageMaker for training and/or inference in production environments.
• Ability to work closely with the hardware layer; familiarity with device-specific compilation and the capacity to debug accelerator-related issues.
• Experience with AWS Trainium or Inferentia (Neuron SDK) is highly desirable.
• Profound PyTorch expertise and a proven ability to quickly adapt to new hardware targets, even if lacking specific Trainium/Inferentia experience.
• Strong foundational knowledge in Python.
• Comfort in a client-facing, production engineering setting.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and growth.
• Flexible work hours and remote work options.
• Collaborative and innovative work environment.
Salve.Inno
General Dynamics Information Technology
Seismic
Get handpicked remote jobs straight to your inbox weekly.