
AI Safety Experts – English, Gujarati
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments on conversational AI models and agents through methods like jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
• Generate reproducible reports, datasets, and case studies on attacks.
• Evaluate AI outputs that involve sensitive subjects, including bias, misinformation, or harmful behaviors.
• Identify vulnerabilities that automated tests may overlook.
• Enhance evaluation coverage and minimize surprises during production.
• Fortify customer AI systems through the creation of reproducible artifacts.
• Proficiency in English and Gujarati at a fluent/native level is essential.
• Strong discernment regarding language and content appropriateness.
• Capability to evaluate AI responses for accuracy, completeness, and relevance, along with the ability to articulate the reasoning behind assessments.
• Meticulous attention to detail, including subtle errors, inconsistencies, and omissions.
• Consistent adherence to established guidelines and quality standards.
• Ability to convey reasoning clearly to both technical and non-technical audiences.
• Flexibility to adapt across various projects, task types, and client needs.
• Preferred areas of expertise include adversarial machine learning, cybersecurity, socio-technical risk, or creative probing.
• Engagement as an independent contractor is required.
• Support for H1-B and STEM OPT candidates is not available.
• Completely remote position.
• Flexible working hours, allowing you to set your own schedule.
• Weekly payments processed through Stripe or Wise.
• Competitive compensation package.
• Opportunity to participate in projects of higher sensitivity, if desired.
• Clear guidelines and wellness resources available for managing sensitive content work.
• Reasonable accommodations can be made upon request.
• Referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.