Machine Learning Engineer, Evals

Posted Jul 28

This is a fully remote position, open to applicants in United States.

📋 Description

• Oversee the complete evaluation pipeline from start to finish and replicate established results during onboarding, collaborating with a senior engineer on the initial task.

• Develop a judge calibration protocol utilizing human-labeled decisions, agreement metrics, drift-zone identification, and comprehensive documentation.

• Enhance benchmarks such as GAIA, τ-Bench, and SWE-bench by incorporating tasks aimed at addressing capability gaps, which include prompts, environments, rubrics, automated grading systems, and quality assurance.

• Evaluate model output failures, classify failure modes, assess their frequency, and suggest modifications to training data, judge prompts, or benchmarks.

• Manage ongoing evaluation workflows, including regression suites, judge-drift dashboards, and red-team evaluations.

• Deliver evaluation tools utilized by researchers.


⛳️ Requirements

• Minimum of 3 years in software engineering, ML engineering, data science, or a research-related position.

• Practical evaluation experience gained through academic courses, internships, side projects, open-source contributions, or employment.

• Familiarity with at least one LLM evaluation framework, such as Harbor or Nemo Evaluator.

• Direct experience with LLMs involving prompting and few-shot design; preferably fine-tuning or retrieval-augmented generation; regular engagement with coding agents.

• Proficient in Python with the ability to write clean, tested, and version-controlled code.

• Comfortable with Git, basic CI/CD practices, Docker, and the Linux command line, including SSH, tmux, and troubleshooting remote jobs.

• Knowledge of fundamental evaluation statistics, encompassing accuracy on imbalanced judges, Cohen's κ, and confidence intervals.

• At least 3 defined evaluation competencies related to judge calibration, failure analysis, agent benchmarks, evaluation dataset design, or reporting on non-determinism and variance.

• Effective communication skills with researchers and engineers.

• Ability to navigate ambiguity, develop plans from incomplete requests, and recognize when to seek assistance.

• Preferred: Experience with RLVR/RLHF pipeline.

• Preferred: Experience in training data curation.

• Preferred: Experience with distributed evaluation orchestration.

• Preferred: Benchmark design from the ground up.

• Preferred: Experience in red teaming and adversarial evaluation.

• Preferred: Familiarity with psychometrics or measurement theory.


🏝️ Benefits

• Comprehensive health and wellness benefits.

• Opportunities for professional development and continuous learning.

• A collaborative and innovative work environment.

• Flexible working hours and remote work options.

People also viewed

Sourcegraph1 day ago

ML Engineer, Agentic Systems

North AmericaFull-timeMachine Learning Engineer$88k – $176k/year
ApplyView job
Quora1 day ago

Senior Machine Learning Engineer, Ads

US flagUnited States OnlyFull-timeMachine Learning Engineer$189.5k – $274.6k/year
ApplyView job
NBCUniversal1 day ago

Staff MLOps Engineer

CA flagCanada OnlyFull-timeMachine Learning Engineer
ApplyView job
Spotify1 day ago

Staff Machine Learning Engineer – Home Surfaces

US flagNew York OnlyFull-timeMachine Learning Engineer
ApplyView job
Fullscript1 day ago

Senior Machine Learning Engineer

CA flagCanada OnlyFull-timeMachine Learning EngineerC$140k – C$160k/year
ApplyView job
Vida Health1 day ago

Principal AI/ML Engineering Lead

US flagUnited States OnlyFull-timeMachine Learning Engineer$250k – $275k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers