AI Evals Engineer – Evaluation Datasets, Ground Truth

Posted Sep 4

This is a fully remote position, open to applicants in United States.

📋 Description

• Break down pipelines and system architecture into assessable modules with clear input-to-expected-output agreements.

• Establish correctness through rubrics, labeling schemas, edge-case policies, and business-relevant tolerances.

• Focus on validation sets that enable the highest level of iteration.

• Retrieve stratified, de-identified production samples that reflect actual traffic and long-tail cases.

• Define the scope and execute human-labeling programs, develop annotation guidelines and calibration sets, assess inter-annotator agreement, and manage vendors or subject-matter experts.

• Create synthetic/oracle-generated ground truth and validate oracle outputs against human-labeled samples.

• Develop programmatic and adversarial labels, templated edge cases, and incident-based backtests.

• Version datasets while tracking lineage, splits, model/prompt exposure, contamination, and leakage.

• Segment datasets by customer type, input category, and complexity; update and retire examples as traffic evolves.

• Adjust automated graders against human gold standards and determine their reliability for gating decisions.

• Provide reports on precision, recall, F1 scores, confusion matrices, calibration, confidence intervals, and sample-size prerequisites.

• Collaborate with eval-harness and CI engineers as well as ML engineers on classifier features.

• Report directly to the Chief AI Officer and collaborate across product and pipeline teams; expected to eventually lead a small eval-engineering team.


⛳️ Requirements

• Candidates must possess authorization to work in the United States without current or future sponsorship.

• Proven experience in building or assessing ML or LLM-powered systems in a production environment, focusing on output quality.

• Experience in constructing evaluation or validation datasets, with the ability to articulate correctness definitions, label sourcing, failure modes, and label quality.

• Proficient understanding of ML validation principles: train/validation/test discipline, stratified sampling, precision/recall trade-offs, class imbalance, calibration, inter-rater agreement, and basic significance testing and power.

• Knowledge of feature engineering sufficient to analyze classifier failures and data needs.

• Strong proficiency in Python and SQL; adept at extracting and reshaping data.

• Hands-on experience with LLM-based systems, including prompting, structured outputs, agent/tool-use harnesses, non-determinism, prompt sensitivity, and evaluator bias.

• Sound judgment regarding when LLM-as-judge is reliable and methods for its validation.

• Capability to read system designs, comprehend business logic, and translate it into labeling schemas.

• Experience managing human-labeling programs from start to finish, including vendor selection, guideline development, QA sampling, and balancing cost/quality (preferred).

• Familiarity with evaluation tooling (preferred).

• Experience with data labeling platforms (preferred).

• Background utilizing a frontier model as a distillation/oracle source (preferred).


🏝️ Benefits

• Complete medical, dental & vision insurance coverage for you; 30% coverage for dependents.

• Competitive salary and significant early-stage equity opportunities.

• Unlimited paid time off (PTO).

• Stipend for hybrid/remote work.

• In-office perks including snacks, beverages, coffee, a ping-pong table, and more.

• Budget allocated for intra-office travel.

• 2–3 annual in-person team meetups.

• Generous budget for labeling vendors and oracle computing resources.

People also viewed

Apheris16 hours ago

AI Program Lead

DE flagGermany OnlyFull-timeArtificial Intelligence
ApplyView job
Mercor17 hours ago

AI Safety Expert – English, Tamil

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor17 hours ago

AI Safety Expert – English, Tamil

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor17 hours ago

AI Safety Expert – English, Kannada

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor17 hours ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
General Dynamics Information Technology17 hours ago

AI/ML Architect

US flagUnited States OnlyFull-timeArtificial Intelligence$128k – $173.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers