
AI Evaluation Guidelines & Rubric Specialist
Posted Sep 12

Posted Sep 12
This is a fully remote position, open to applicants in New York.
β’ Translate program requirements into clear, structured, and actionable instructions for human evaluators.
β’ Develop guidelines relevant to standard scenarios as well as complex edge cases.
β’ Define terminology, rating criteria, decision rules, exceptions, and escalation pathways.
β’ Ensure that instructions are accessible to raters while maintaining domain-specific precision.
β’ Design detailed scoring rubrics for evaluating generative AI outputs.
β’ Establish measurable criteria for correctness, relevance, reasoning quality, completeness, and adherence to instructions.
β’ Create examples and counterexamples that illustrate varying performance levels.
β’ Align evaluation frameworks with program objectives and quality standards.
β’ Review draft specifications to identify ambiguity, contradiction, missing information, and inconsistent terminology.
β’ Identify instructions that could lead to conflicting interpretations among raters.
β’ Revise guideline sets to ensure reliable application with minimal escalation.
β’ Document improvements to written requirements and evaluation instructions before and after revisions.
β’ Transform specialist-domain specifications into guidance suitable for raters across finance, retail, insurance, legal, sports, and other fields.
β’ Collaborate with subject matter experts to ensure accurate domain-specific terminology and professional judgment.
β’ Maintain technical nuance while ensuring instructions are comprehensible to non-specialist evaluators.
β’ Ensure a consistent structure and quality across domain-specific guideline sets.
β’ Participate in onboarding, calibration, documentation reviews, and ongoing refinement of guidelines.
β’ A minimum of 3 years of professional experience in linguistics, instructional design, technical writing, content design, or a closely related field.
β’ Direct experience in developing or refining guidelines and rubrics for human evaluators in generative AI, RLHF, or model-assessment programs.
β’ Proven experience in developing rater guidelines or rubrics specifically for generative AI or RLHF programs is essential.
β’ Demonstrated capability to resolve ambiguity and contradictions within complex written specifications.
β’ Experience in translating specialist requirements into clear and actionable instructions.
β’ Ability to work effectively across diverse subject-matter domains.
β’ A portfolio or specific examples that showcase measurable improvements to guidelines, rubrics, or instructional materials.
β’ Evidence of professional growth and increasing levels of responsibility.
β’ Reliable availability for a minimum of 35 hours per week during weekdays.
β’ Equivalent professional experience in AI evaluation, technical documentation, or guideline development may also be accepted.
β’ Applicants should be ready to present concrete examples of improvements made to guidelines or specifications.
β’ Immediate availability is preferred.
β’ Experience in supporting large language model evaluation, reinforcement learning from human feedback, or AI training-data programs.
β’ Familiarity with annotation platforms, human-feedback workflows, and rater calibration processes.
β’ Experience in developing domain-specific guidance for finance, insurance, retail, legal, sports, or similar fields.
β’ Knowledge of controlled language, information architecture, taxonomy design, or content governance.
β’ Experience conducting usability tests for guidelines or analyzing inter-rater consistency.
β’ Familiarity with version control, documentation systems, and structured authoring tools.
β’ Previous collaboration with researchers, program managers, engineers, and subject matter experts.
β’ Full-time remote engagement.
β’ Competitive hourly compensation ranging from $40β$60 per hour, depending on expertise and project scope.
β’ W-2 contingent employment arrangement.
β’ Remote work opportunity from anywhere within the United States.
β’ Expected commitment of at least 35 hours per week during weekdays.
β’ Opportunities for onboarding and calibration.
β’ Project scope and duration may be adjusted based on program requirements and performance.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.