
Staff ML Engineer – AWS Trainium, SageMaker
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Canada.
• Train and manage models using Amazon SageMaker, leveraging AWS Trainium as the core computing resource.
• Develop and refine PyTorch training scripts for Trainium, considering NeuronCore architecture, compiler functionality, memory management, and throughput optimization.
• Identify and resolve hardware-related training issues, differentiating between data or code challenges and compiler or device-level complications.
• Convert Trainium job specifications into efficient, cost-effective end-to-end training workflows.
• Optimize distributed training sessions for performance and cost-effectiveness on SageMaker training infrastructure.
• Collaborate closely with clients and internal engineering teams to define and deliver production-grade training workloads.
• Extensive hands-on experience with PyTorch, particularly in distributed or multi-device training scenarios.
• Practical experience utilizing Amazon SageMaker for training and/or inference processes.
• Ability to work near the hardware level, including device-specific compilation and debugging at the accelerator level.
• Experience with AWS Trainium or Inferentia (Neuron SDK) is highly advantageous.
• Strong background in PyTorch and a proven ability to quickly adapt to new hardware targets may compensate for lack of Trainium experience.
• Solid foundational knowledge of Python.
• Experience in a client-facing production engineering environment is essential.
• Opportunities for professional growth and development.
• Collaborative and innovative work environment.
• Competitive compensation and benefits package.
Miratech
First American
Humana
AAA
Get handpicked remote jobs straight to your inbox weekly.