
Machine Learning Engineer, Evals
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in United States.
• Oversee the complete evaluation pipeline from start to finish and replicate established results during onboarding, collaborating with a senior engineer on the initial task.
• Develop a judge calibration protocol utilizing human-labeled decisions, agreement metrics, drift-zone identification, and comprehensive documentation.
• Enhance benchmarks such as GAIA, τ-Bench, and SWE-bench by incorporating tasks aimed at addressing capability gaps, which include prompts, environments, rubrics, automated grading systems, and quality assurance.
• Evaluate model output failures, classify failure modes, assess their frequency, and suggest modifications to training data, judge prompts, or benchmarks.
• Manage ongoing evaluation workflows, including regression suites, judge-drift dashboards, and red-team evaluations.
• Deliver evaluation tools utilized by researchers.
• Minimum of 3 years in software engineering, ML engineering, data science, or a research-related position.
• Practical evaluation experience gained through academic courses, internships, side projects, open-source contributions, or employment.
• Familiarity with at least one LLM evaluation framework, such as Harbor or Nemo Evaluator.
• Direct experience with LLMs involving prompting and few-shot design; preferably fine-tuning or retrieval-augmented generation; regular engagement with coding agents.
• Proficient in Python with the ability to write clean, tested, and version-controlled code.
• Comfortable with Git, basic CI/CD practices, Docker, and the Linux command line, including SSH, tmux, and troubleshooting remote jobs.
• Knowledge of fundamental evaluation statistics, encompassing accuracy on imbalanced judges, Cohen's κ, and confidence intervals.
• At least 3 defined evaluation competencies related to judge calibration, failure analysis, agent benchmarks, evaluation dataset design, or reporting on non-determinism and variance.
• Effective communication skills with researchers and engineers.
• Ability to navigate ambiguity, develop plans from incomplete requests, and recognize when to seek assistance.
• Preferred: Experience with RLVR/RLHF pipeline.
• Preferred: Experience in training data curation.
• Preferred: Experience with distributed evaluation orchestration.
• Preferred: Benchmark design from the ground up.
• Preferred: Experience in red teaming and adversarial evaluation.
• Preferred: Familiarity with psychometrics or measurement theory.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and continuous learning.
• A collaborative and innovative work environment.
• Flexible working hours and remote work options.
Sourcegraph
Quora
NBCUniversal
Spotify
Get handpicked remote jobs straight to your inbox weekly.