
AI Safety Expert, English – Kannada
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments of conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
• Generate human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
• Create reproducible reports, datasets, and attack cases for clients.
• Evaluate AI outputs related to sensitive subjects, including bias, misinformation, and harmful behaviors.
• Investigate AI systems to detect vulnerabilities that automated tests may overlook.
• Enhance evaluation coverage and minimize unexpected outcomes in production.
• Contribute to the training and improvement of frontier AI models for Mercor’s AI lab and enterprise clients.
• Native fluency in both English and Kannada.
• Strong discernment regarding language and content, including evaluating the accuracy, completeness, and appropriateness of AI responses.
• Capability to clearly articulate reasoning to both technical and non-technical audiences.
• Meticulous attention to detail, including identifying subtle errors, inconsistencies, and omissions.
• Consistent adherence to taxonomies, benchmarks, playbooks, guidelines, and quality standards.
• Flexibility across various projects, task types, and client needs.
• Status as an independent contractor.
• Must not require H1-B or STEM OPT sponsorship; H1-B and STEM OPT candidates cannot be accommodated.
• Experience in adversarial ML, cybersecurity, socio-technical risk, conversational AI testing, psychology, acting, or unconventional adversarial writing is a plus.
• Fully remote position that allows for flexible scheduling.
• Weekly payments through Stripe or Wise based on services provided.
• Option to participate in higher-sensitivity projects.
• Access to clear guidelines and wellness resources for projects involving sensitive content.
• Competitive compensation.
• Opportunity to gain experience in human data-driven AI red teaming.
• Collaborate with leading researchers in the field.
• Reasonable accommodations available upon request.
• Earn referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.