
Staff Machine Learning Engineer, Ads ML Efficiency
Posted 19 hours ago

Posted 19 hours ago
This is a fully remote position, open to applicants in United States, +1 more country.
• Design and develop systems that enhance the efficiency of machine learning training and inference processes.
• Create tools that assist ML engineers in debugging, profiling, optimizing, and monitoring model performance.
• Enhance GPU and overall resource utilization through effective scheduling, resource management, caching, and optimization of workloads.
• Collaborate with ML researchers and product teams to pinpoint bottlenecks and implement performance enhancements.
• Construct benchmarking frameworks and performance dashboards for both training and serving systems.
• Optimize distributed training infrastructure, data pipelines, and architectures for model serving.
• Lead cross-functional projects that boost the productivity of Reddit's ML engineers.
• Advocate for technical strategies that enhance the scalability, reliability, and cost-effectiveness of the ML platform.
• Empower ML engineers to transition from ideas to experiments more rapidly.
• Reduce training and inference costs while enhancing performance and ensuring or improving model quality.
• Increase GPU utilization and overall cluster efficiency.
• Enhance platform reliability as ML workloads grow.
• Bachelor's, Master's, or PhD degree in Computer Science or a related discipline.
• Over 5 years of experience in software engineering.
• Strong expertise in Python programming.
• Proficiency in at least one systems programming language (preferably Go, C++, Rust, or Java).
• Experience in building distributed systems at scale.
• Familiarity with machine learning infrastructure, training systems, or model serving platforms.
• In-depth knowledge of performance engineering and systems optimization techniques.
• Excellent debugging and profiling capabilities.
• Preferred: experience with large-scale recommendation, ranking, generative AI, or foundational model systems.
• Preferred: experience with distributed training frameworks such as PyTorch Distributed, Ray, TensorFlow, or Spark.
• Preferred: understanding of GPU architectures and performance analysis tools.
• Preferred: experience in optimizing cloud infrastructure costs across extensive ML workloads.
• Preferred: contributions to internal platforms utilized by multiple ML teams.
• Preferred: experience in developing real-time ML inference applications.
• Global benefit programs tailored to your lifestyle, including workspace, professional development, and caregiving support.
• Family Planning Support.
• Gender-Affirming Care.
• Mental Health & Coaching Benefits.
• Group Personal Pension Scheme with Employer match.
• Private Medical and Dental Scheme.
• Income Replacement Programs.
• Bike to Work scheme.
• Flexible Vacation & Paid Volunteer Time Off.
• Generous Paid Parental Leave.
• Equity in the form of restricted stock units.
• Medical, dental, and vision insurance.
• 401(k) program with employer match.
• Generous time off for vacation and parental leave.
Defcon AI
Stack AV
Solventum
MUTT DATA
Get handpicked remote jobs straight to your inbox weekly.