
AI Evals Engineer – Evaluation Datasets, Ground Truth
Posted Sep 4

Posted Sep 4
This is a fully remote position, open to applicants in United States.
• Break down pipelines and system architecture into assessable modules with clear input-to-expected-output agreements.
• Establish correctness through rubrics, labeling schemas, edge-case policies, and business-relevant tolerances.
• Focus on validation sets that enable the highest level of iteration.
• Retrieve stratified, de-identified production samples that reflect actual traffic and long-tail cases.
• Define the scope and execute human-labeling programs, develop annotation guidelines and calibration sets, assess inter-annotator agreement, and manage vendors or subject-matter experts.
• Create synthetic/oracle-generated ground truth and validate oracle outputs against human-labeled samples.
• Develop programmatic and adversarial labels, templated edge cases, and incident-based backtests.
• Version datasets while tracking lineage, splits, model/prompt exposure, contamination, and leakage.
• Segment datasets by customer type, input category, and complexity; update and retire examples as traffic evolves.
• Adjust automated graders against human gold standards and determine their reliability for gating decisions.
• Provide reports on precision, recall, F1 scores, confusion matrices, calibration, confidence intervals, and sample-size prerequisites.
• Collaborate with eval-harness and CI engineers as well as ML engineers on classifier features.
• Report directly to the Chief AI Officer and collaborate across product and pipeline teams; expected to eventually lead a small eval-engineering team.
• Candidates must possess authorization to work in the United States without current or future sponsorship.
• Proven experience in building or assessing ML or LLM-powered systems in a production environment, focusing on output quality.
• Experience in constructing evaluation or validation datasets, with the ability to articulate correctness definitions, label sourcing, failure modes, and label quality.
• Proficient understanding of ML validation principles: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, and basic significance testing and power.
• Knowledge of feature engineering sufficient to analyze classifier failures and data needs.
• Strong proficiency in Python and SQL; adept at extracting and reshaping data.
• Hands-on experience with LLM-based systems, including prompting, structured outputs, agent/tool-use harnesses, non-determinism, prompt sensitivity, and evaluator bias.
• Sound judgment regarding when LLM-as-judge is reliable and methods for its validation.
• Capability to read system designs, comprehend business logic, and translate it into labeling schemas.
• Experience managing human-labeling programs from start to finish, including vendor selection, guideline development, QA sampling, and balancing cost/quality (preferred).
• Familiarity with evaluation tooling (preferred).
• Experience with data labeling platforms (preferred).
• Background utilizing a frontier model as a distillation/oracle source (preferred).
• Complete medical, dental & vision insurance coverage for you; 30% coverage for dependents.
• Competitive salary and significant early-stage equity opportunities.
• Unlimited paid time off (PTO).
• Stipend for hybrid/remote work.
• In-office perks including snacks, beverages, coffee, a ping-pong table, and more.
• Budget allocated for intra-office travel.
• 2–3 annual in-person team meetups.
• Generous budget for labeling vendors and oracle computing resources.
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.