
AI Safety Expert – English, Telugu
Posted Sep 19

Posted Sep 19
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments of conversational AI models and agents through techniques such as jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
• Create human data by documenting failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
• Generate reproducible reports, datasets, and attack scenarios for clients.
• Analyze AI outputs related to sensitive subjects like bias, misinformation, and harmful behaviors.
• Identify vulnerabilities that automated testing may overlook.
• Broaden evaluation coverage and minimize production surprises.
• Collaborate on projects focused on training and enhancing AI systems for Mercor’s AI lab and its enterprise clients.
• Native or fluent proficiency in English and Telugu is essential.
• Strong discernment regarding language and content.
• Capability to evaluate the accuracy, completeness, and appropriateness of AI responses, along with the ability to articulate reasoning.
• Keen attention to detail in identifying subtle errors, inconsistencies, and gaps.
• Consistent adherence to guidelines and quality standards.
• Ability to communicate reasoning effectively to both technical and non-technical audiences.
• Flexibility across various projects, task types, and clients.
• Status as an independent contractor.
• H1-B and STEM OPT candidates are not eligible.
• Preferred: experience in adversarial ML, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
• Preferred: experience in cybersecurity, including penetration testing, exploit development, or reverse engineering.
• Preferred: experience in socio-technical risk management, encompassing harassment/disinformation probing, abuse analysis, or conversational AI testing.
• Preferred: background in psychology, acting, or writing to foster unconventional adversarial thinking.
• Fully remote position.
• Flexible working hours, allowing you to set your own schedule.
• Weekly payments facilitated through Stripe or Wise.
• Participation in higher-sensitivity projects is optional.
• Comprehensive guidelines and wellness resources available for sensitive-content work.
• Project durations may be extended, shortened, or concluded early based on requirements and performance.
• Competitive compensation.
• Referral bonus of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.