Senior Deep Learning Engineer, Accuracy Evaluation

Posted Aug 21

This is a fully remote position, open to applicants in Poland, +4 more countries.

📋 Description

• Design and construct decision-grade evaluation environments for NVIDIA's advanced models, including those focused on reasoning, multimodal, long-context, and agentic systems.

• Generate auditable accuracy signals that are essential for every major model release.

• Investigate and develop innovative evaluation methodologies for new model families and capability domains.

• Create and maintain evaluation infrastructure and pipelines, encompassing benchmark environments, regression CI systems, and statistical analysis tools.

• Collaborate with model research, training, and customer teams to convert evaluation signals into release decisions, guide training iterations, and enhance competitive positioning.

• Work together with various teams at NVIDIA to realize flagship models from the community and partners.

• Deliver advanced models with rapid inference capabilities utilizing enterprise-grade GPU clusters.


⛳️ Requirements

• BS, MS, or PhD in Computer Science, Machine Learning, Statistics, or a related discipline.

• Over 6 years of practical experience with large language models (LLMs).

• Proficiency in designing and executing evaluations for large language models or multimodal AI systems.

• Experience in agentic, multi-turn, or reasoning-intensive settings.

• Strong foundation in statistics, including experimental design, significance testing, and regression analysis.

• Demonstrated experience in developing evaluation infrastructure, such as pipelines, benchmark harnesses, and reproducible CI systems.

• Capability to discern signal from noise in benchmark results on a large scale.

• Ability to convert quantitative evaluation results into actionable insights for researchers, product teams, and senior leadership.

• Expertise in deep learning and AI evaluation.

• Familiarity with open-source evaluation frameworks.

• Experience in designing evaluations for agentic systems.

• A proven track record of publishing or contributing to evaluation research, benchmark design, methodology papers, or reproducibility analyses.

• Experience in measuring model accuracy in low-precision inference settings, including FP8, INT4, and quantization-aware approaches.

• Comfort with managing large-scale workloads on HPC/Slurm clusters.

• Experience with reproducible experiment management using MLflow or W&B.

• Proven ability to optimize compute costs across numerous benchmark runs.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health and wellness benefits.

• Opportunities for professional development and growth.

• Flexible work hours and remote work options.

• Access to cutting-edge technology and resources.

• Collaborative and inclusive company culture.

People also viewed

Slate Auto1 day ago

AI/ML Engineer

US flagMichigan OnlyFull-timeMachine Learning Engineer$123.3k – $185k/year
ApplyView job
Pinterest1 day ago

Senior Machine Learning Engineer

US flagCalifornia OnlyFull-timeMachine Learning Engineer$277k – $332k/year
ApplyView job
Pinterest1 day ago

Senior Machine Learning Engineer

US flagCalifornia OnlyFull-timeMachine Learning Engineer$246.9k – $332k/year
ApplyView job
Pinterest1 day ago

Machine Learning Engineer II

US flagCalifornia OnlyFull-timeMachine Learning Engineer$215.3k – $286k/year
ApplyView job
Compass1 day ago

Senior Machine Learning Engineer

BR flagBrazil OnlyFull-timeMachine Learning Engineer
ApplyView job
Netflix1 day ago

Machine Learning Scientist 5 – Ads Demand Science

US flagUnited States OnlyFull-timeMachine Learning Engineer$466k – $750k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers