
AI Safety Expert – English, Kannada
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create human-generated data by annotating failures, identifying vulnerabilities, and highlighting systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure thorough and consistent testing.
• Generate reproducible reports, datasets, and attack scenarios for clients.
• Evaluate AI outputs related to sensitive subjects including bias, misinformation, and harmful behaviors.
• Identify vulnerabilities that automated tests may overlook.
• Broaden evaluation coverage and minimize unexpected issues in production.
• Enhance customer AI systems through adversarial testing techniques.
• Engage in projects aimed at training and improving cutting-edge AI systems.
• Proficiency in both English and Kannada is essential.
• Strong discernment regarding language and content quality.
• Capable of evaluating whether AI responses are accurate, comprehensive, and suitable, along with the ability to articulate reasoning.
• Meticulous attention to detail regarding subtle errors, inconsistencies, and omissions.
• Consistent adherence to taxonomies, benchmarks, playbooks, guidelines, and quality standards.
• Ability to communicate reasoning clearly to both technical and non-technical audiences.
• Flexibility to adapt across various projects, tasks, and client needs.
• Engagement as an independent contractor.
• H1-B and STEM OPT candidates are not eligible.
• Preferred expertise includes adversarial ML, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation probing, abuse analysis, conversational AI testing, psychology, acting, and writing.
• Fully remote position.
• Flexible scheduling to suit individual preferences.
• Weekly payments processed through Stripe or Wise.
• Optional participation in higher-sensitivity projects.
• Clear guidelines and wellness resources for managing sensitive content.
• Competitive compensation.
• Opportunity to collaborate with leading researchers in the field.
• Referral bonuses of up to $90 for each successful referral, subject to limitations.
• Reasonable accommodations available upon request.
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.