Remotery

Senior Deep Learning Engineer, Accuracy Evaluation

Posted 6 hours ago

This is a fully remote position, open to applicants in Poland, +4 more countries.

📋 Description

• Design and construct decision-grade evaluation environments for NVIDIA's advanced models, including those focused on reasoning, multimodal, long-context, and agentic systems.

• Generate auditable accuracy signals that are essential for every major model release.

• Investigate and develop innovative evaluation methodologies for new model families and capability domains.

• Create and maintain evaluation infrastructure and pipelines, encompassing benchmark environments, regression CI systems, and statistical analysis tools.

• Collaborate with model research, training, and customer teams to convert evaluation signals into release decisions, guide training iterations, and enhance competitive positioning.

• Work together with various teams at NVIDIA to realize flagship models from the community and partners.

• Deliver advanced models with rapid inference capabilities utilizing enterprise-grade GPU clusters.


⛳️ Requirements

• BS, MS, or PhD in Computer Science, Machine Learning, Statistics, or a related discipline.

• Over 6 years of practical experience with large language models (LLMs).

• Proficiency in designing and executing evaluations for large language models or multimodal AI systems.

• Experience in agentic, multi-turn, or reasoning-intensive settings.

• Strong foundation in statistics, including experimental design, significance testing, and regression analysis.

• Demonstrated experience in developing evaluation infrastructure, such as pipelines, benchmark harnesses, and reproducible CI systems.

• Capability to discern signal from noise in benchmark results on a large scale.

• Ability to convert quantitative evaluation results into actionable insights for researchers, product teams, and senior leadership.

• Expertise in deep learning and AI evaluation.

• Familiarity with open-source evaluation frameworks.

• Experience in designing evaluations for agentic systems.

• A proven track record of publishing or contributing to evaluation research, benchmark design, methodology papers, or reproducibility analyses.

• Experience in measuring model accuracy in low-precision inference settings, including FP8, INT4, and quantization-aware approaches.

• Comfort with managing large-scale workloads on HPC/Slurm clusters.

• Experience with reproducible experiment management using MLflow or W&B.

• Proven ability to optimize compute costs across numerous benchmark runs.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health and wellness benefits.

• Opportunities for professional development and growth.

• Flexible work hours and remote work options.

• Access to cutting-edge technology and resources.

• Collaborative and inclusive company culture.

People also viewed

SentiLink7 hours ago

Head of Applied Machine Learning – Application Fraud

US flagUnited States OnlyFull-timeMachine Learning Engineer$210k – $260k/year
ApplyView job
SentiLink7 hours ago

Applied Machine Learning Manager – Application Fraud

US flagUnited States OnlyFull-timeMachine Learning Engineer$200k – $250k/year
ApplyView job
Leega7 hours ago

Senior AI (Generative, MLOps)

BR flagBrazil OnlyFreelanceMachine Learning Engineer
ApplyView job
Solidgate10 hours ago

Senior Machine Learning Engineer

UA flagUkraine OnlyFull-timeMachine Learning Engineer
ApplyView job
FIT:MATCH.ai11 hours ago

Machine Learning Engineer, 3D Vision

US flagCalifornia, +3 more statesFull-timeMachine Learning Engineer
ApplyView job
Instacart11 hours ago

Machine Learning Engineer II – Ads, Response Prediction

CA flagCanada OnlyFull-timeMachine Learning EngineerC$154k – C$162.5k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers