
Senior Deep Learning Engineer, Accuracy Evaluation
Posted 6 hours ago

Posted 6 hours ago
This is a fully remote position, open to applicants in Poland, +4 more countries.
• Design and construct decision-grade evaluation environments for NVIDIA's advanced models, including those focused on reasoning, multimodal, long-context, and agentic systems.
• Generate auditable accuracy signals that are essential for every major model release.
• Investigate and develop innovative evaluation methodologies for new model families and capability domains.
• Create and maintain evaluation infrastructure and pipelines, encompassing benchmark environments, regression CI systems, and statistical analysis tools.
• Collaborate with model research, training, and customer teams to convert evaluation signals into release decisions, guide training iterations, and enhance competitive positioning.
• Work together with various teams at NVIDIA to realize flagship models from the community and partners.
• Deliver advanced models with rapid inference capabilities utilizing enterprise-grade GPU clusters.
• BS, MS, or PhD in Computer Science, Machine Learning, Statistics, or a related discipline.
• Over 6 years of practical experience with large language models (LLMs).
• Proficiency in designing and executing evaluations for large language models or multimodal AI systems.
• Experience in agentic, multi-turn, or reasoning-intensive settings.
• Strong foundation in statistics, including experimental design, significance testing, and regression analysis.
• Demonstrated experience in developing evaluation infrastructure, such as pipelines, benchmark harnesses, and reproducible CI systems.
• Capability to discern signal from noise in benchmark results on a large scale.
• Ability to convert quantitative evaluation results into actionable insights for researchers, product teams, and senior leadership.
• Expertise in deep learning and AI evaluation.
• Familiarity with open-source evaluation frameworks.
• Experience in designing evaluations for agentic systems.
• A proven track record of publishing or contributing to evaluation research, benchmark design, methodology papers, or reproducibility analyses.
• Experience in measuring model accuracy in low-precision inference settings, including FP8, INT4, and quantization-aware approaches.
• Comfort with managing large-scale workloads on HPC/Slurm clusters.
• Experience with reproducible experiment management using MLflow or W&B.
• Proven ability to optimize compute costs across numerous benchmark runs.
• Competitive salary and performance-based bonuses.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and growth.
• Flexible work hours and remote work options.
• Access to cutting-edge technology and resources.
• Collaborative and inclusive company culture.
SentiLink
SentiLink
Leega
Solidgate
Get handpicked remote jobs straight to your inbox weekly.