AI Evaluation Guidelines & Rubric Specialist

at24-MAGRemoteUS flagNew YorkFull-timeArtificial IntelligenceMid-levelSenior$40 – $60/hour

Posted Sep 12

This is a fully remote position, open to applicants in New York.

πŸ“‹ Description

β€’ Translate program requirements into clear, structured, and actionable instructions for human evaluators.

β€’ Develop guidelines relevant to standard scenarios as well as complex edge cases.

β€’ Define terminology, rating criteria, decision rules, exceptions, and escalation pathways.

β€’ Ensure that instructions are accessible to raters while maintaining domain-specific precision.

β€’ Design detailed scoring rubrics for evaluating generative AI outputs.

β€’ Establish measurable criteria for correctness, relevance, reasoning quality, completeness, and adherence to instructions.

β€’ Create examples and counterexamples that illustrate varying performance levels.

β€’ Align evaluation frameworks with program objectives and quality standards.

β€’ Review draft specifications to identify ambiguity, contradiction, missing information, and inconsistent terminology.

β€’ Identify instructions that could lead to conflicting interpretations among raters.

β€’ Revise guideline sets to ensure reliable application with minimal escalation.

β€’ Document improvements to written requirements and evaluation instructions before and after revisions.

β€’ Transform specialist-domain specifications into guidance suitable for raters across finance, retail, insurance, legal, sports, and other fields.

β€’ Collaborate with subject matter experts to ensure accurate domain-specific terminology and professional judgment.

β€’ Maintain technical nuance while ensuring instructions are comprehensible to non-specialist evaluators.

β€’ Ensure a consistent structure and quality across domain-specific guideline sets.

β€’ Participate in onboarding, calibration, documentation reviews, and ongoing refinement of guidelines.


⛳️ Requirements

β€’ A minimum of 3 years of professional experience in linguistics, instructional design, technical writing, content design, or a closely related field.

β€’ Direct experience in developing or refining guidelines and rubrics for human evaluators in generative AI, RLHF, or model-assessment programs.

β€’ Proven experience in developing rater guidelines or rubrics specifically for generative AI or RLHF programs is essential.

β€’ Demonstrated capability to resolve ambiguity and contradictions within complex written specifications.

β€’ Experience in translating specialist requirements into clear and actionable instructions.

β€’ Ability to work effectively across diverse subject-matter domains.

β€’ A portfolio or specific examples that showcase measurable improvements to guidelines, rubrics, or instructional materials.

β€’ Evidence of professional growth and increasing levels of responsibility.

β€’ Reliable availability for a minimum of 35 hours per week during weekdays.

β€’ Equivalent professional experience in AI evaluation, technical documentation, or guideline development may also be accepted.

β€’ Applicants should be ready to present concrete examples of improvements made to guidelines or specifications.

β€’ Immediate availability is preferred.

β€’ Experience in supporting large language model evaluation, reinforcement learning from human feedback, or AI training-data programs.

β€’ Familiarity with annotation platforms, human-feedback workflows, and rater calibration processes.

β€’ Experience in developing domain-specific guidance for finance, insurance, retail, legal, sports, or similar fields.

β€’ Knowledge of controlled language, information architecture, taxonomy design, or content governance.

β€’ Experience conducting usability tests for guidelines or analyzing inter-rater consistency.

β€’ Familiarity with version control, documentation systems, and structured authoring tools.

β€’ Previous collaboration with researchers, program managers, engineers, and subject matter experts.


🏝️ Benefits

β€’ Full-time remote engagement.

β€’ Competitive hourly compensation ranging from $40–$60 per hour, depending on expertise and project scope.

β€’ W-2 contingent employment arrangement.

β€’ Remote work opportunity from anywhere within the United States.

β€’ Expected commitment of at least 35 hours per week during weekdays.

β€’ Opportunities for onboarding and calibration.

β€’ Project scope and duration may be adjusted based on program requirements and performance.

People also viewed

Mercor13 hours ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor14 hours ago

AI Safety Expert – English, Telugu

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor14 hours ago

AI Safety Experts, English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor23 hours ago

AI Safety Expert, English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor23 hours ago

AI Safety Experts – English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor23 hours ago

AI Safety Expert – English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers