
AI Safety Expert, English, Gujarati
Posted Sep 19

Posted Sep 19
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments of conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
• Generate human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistency in testing.
• Create reproducible reports, datasets, and attack scenarios.
• Evaluate AI outputs related to sensitive subjects, including bias, misinformation, and harmful behaviors.
• Identify vulnerabilities overlooked by automated testing.
• Broaden evaluation coverage and enhance customer AI systems.
• Native fluency in both English and Gujarati is essential.
• Strong judgment regarding language and content; capable of determining whether AI responses are accurate, complete, and suitable.
• Proficient in spotting subtle errors, inconsistencies, and gaps.
• Ability to consistently adhere to taxonomies, benchmarks, playbooks, guidelines, and quality standards.
• Skillful in articulating reasoning clearly to both technical and non-technical audiences.
• Versatile across various projects, task types, and client needs.
• Status as an independent contractor.
• Must not require H1-B or STEM OPT sponsorship.
• Preferred areas of expertise include adversarial ML, cybersecurity, socio-technical risk, or creative probing.
• Completely remote position.
• Flexible working hours / manage your own schedule.
• Weekly payments through Stripe or Wise.
• Optional participation in higher-sensitivity projects.
• Clear guidelines and wellness resources available for sensitive-content work.
• Reasonable accommodations provided upon request.
• Competitive compensation.
• Referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.