
Data Annotation Specialist – Generalist
Posted Jul 24

Posted Jul 24
This is a fully remote position, open to applicants in Canada.
• Assess and rank model outputs: Perform preference and comparison tasks to determine which responses align best with project guidelines regarding accuracy, helpfulness, tone, and safety, while providing clear justifications for your conclusions.
• Stress-test and analyze models: Investigate models adversarially to identify failure modes, unsafe behaviors, and capability gaps, documenting reproducible instances for engineering and research teams to address.
• Develop datasets: Create high-quality prompts, responses, and examples to construct training and evaluation datasets, adhering to detailed specifications and refining both machine-generated and human-generated outputs to meet standards.
• Establish and implement rubrics and taxonomies: Assist in designing grading criteria and rubrics, then consistently apply them to yield structured, high-quality annotations across various task types.
• Annotate and rectify multimodal data: Label, audit, and correct inaccuracies across text, images, and structured data, ensuring a high level of data integrity and precision.
• Calibrate and uphold consistency: Engage in calibration exercises and inter-annotator agreement checks to establish alignment on standards, and identify ambiguous or unexplored edge cases instead of guessing, as a single misjudgment replicated at scale can undermine a model.
• Adapt to experimental tasks: Embrace new and evolving task types as project requirements change, applying sound judgment in areas where guidelines are still in development.
• Report on model performance: Highlight and communicate quality and performance trends in model and agent behavior, providing cross-functional partners with clear, well-supported feedback on model successes, failures, and degradation.
• Over 1 year of experience in AI data annotation, LLM evaluation, content moderation, research, or a similar analytical role, with familiarity in quality assurance and/or preference ranking.
• Experience applying detailed guidelines to complex and often ambiguous content, demonstrating strong contextual and sociocultural judgment, sensitivity to nuance, tone, and register, along with the ability to reason effectively in situations lacking a single correct answer.
• Comfort with ambiguity: a readiness to highlight unclear or uncovered edge cases rather than making guesses, and to work productively on novel, experimental tasks with evolving definitions.
• A keen, inquisitive eye for inconsistencies, subtle errors, and model failure modes, including an instinct to challenge models adversarially and identify their weaknesses. Familiarity with the behavior of large language models, such as hallucination, sycophancy, and instruction-following gaps, is an advantage.
• Excellent command of written English and strong reading comprehension skills, with the capability to clearly justify your evaluations, including the rationale behind why an output is deemed correct or incorrect, high-quality or low-quality, as well as crafting clear prompts and examples that clarify your reasoning. Bonus points for fluency in another language!
• Strong attention to detail and commitment to accuracy, with the capability to maintain consistency across high-volume and repetitive tasks.
• Comfortable working with annotation platforms and structured formats such as JSON, CSV/TSV, Markdown, XML, and YAML.
• Strong performance in a remote work setting, including effective time management, ease of using new tools, and the ability to work independently within a global, asynchronous team.
• Bring Your Own Device (BYOD) 💻
Lee, Brock e Camargo Advogados
AppliedVR
CorroHealth
Get handpicked remote jobs straight to your inbox weekly.