
AI Safety Expert, English & Marathi
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Generate human data by annotating instances of failure, classifying vulnerabilities, and identifying systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing methodologies.
β’ Create reproducible reports, datasets, and attack scenarios for clients.
β’ Investigate AI outputs for issues related to bias, misinformation, and harmful behaviors.
β’ Discover vulnerabilities that automated testing may overlook.
β’ Enhance evaluation coverage while minimizing unexpected issues in production.
β’ Proficiency in English and Marathi is mandatory.
β’ Strong discernment regarding language and content, especially in assessing the accuracy, completeness, and appropriateness of AI-generated responses.
β’ Ability to articulate reasoning clearly to both technical and non-technical stakeholders.
β’ Meticulous attention to detail, including the identification of subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to established guidelines and quality standards.
β’ Flexibility across various projects, task types, and client needs.
β’ Status as an independent contractor is required.
β’ H1-B and STEM OPT candidates will not be considered.
β’ Fully remote position.
β’ Flexible working hours / ability to set your own schedule.
β’ Weekly payments through Stripe or Wise.
β’ Access to wellness resources for projects with higher sensitivity.
β’ Clear guidelines provided prior to engaging with sensitive content.
β’ Project durations may be adjusted based on needs and performance.
β’ Competitive compensation.
β’ Reasonable accommodations available upon request.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.