
AI Safety Experts – English, Telugu
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments on conversational AI models and agents through various techniques such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
• Generate reproducible reports, datasets, and attack scenarios for clients.
• Evaluate AI outputs related to sensitive subjects, including bias, misinformation, and harmful behaviors.
• Identify vulnerabilities that automated testing may overlook.
• Broaden evaluation coverage and minimize unexpected issues in production.
• Enhance customer AI systems through adversarial testing methodologies.
• Native proficiency in both English and Telugu.
• Strong judgment regarding language and content; capable of assessing the accuracy, completeness, and appropriateness of AI responses.
• Ability to articulate reasoning clearly to both technical and non-technical audiences.
• Meticulous attention to detail concerning subtle errors, inconsistencies, and gaps.
• Consistent adherence to guidelines and quality standards.
• Flexibility to adapt across different projects, task types, and client needs.
• Engagement as an independent contractor.
• Must not qualify as an H1-B or STEM OPT candidate.
• Desired: experience in adversarial ML, including jailbreak datasets, prompt injections, RLHF/DPO attacks, or model extraction.
• Desired: background in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
• Desired: experience in socio-technical risk, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
• Desired: experience in psychology, acting, or writing that facilitates unconventional adversarial thinking.
• Fully remote position.
• Flexible work schedule; manage your own hours.
• Receive weekly payments via Stripe or Wise.
• Project durations may vary based on needs and performance.
• Participation in higher-sensitivity projects is optional.
• Access to clear guidelines and wellness resources for handling sensitive content.
• Competitive compensation.
• Opportunity to collaborate with leading researchers.
• Reasonable accommodations available upon request.
• Referral bonuses of up to $90 for each successful referral (referral limits apply).
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.