
AI Safety Experts, English – Gujarati
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team evaluations on conversational AI models and agents through jailbreak techniques, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Generate human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
• Create reproducible reports, datasets, and attack cases for client use.
• Analyze AI outputs related to sensitive subjects such as bias, misinformation, and harmful behavior.
• Expose vulnerabilities that automated tests may overlook.
• Enhance evaluation coverage and minimize unexpected issues during production.
• Collaborate on initiatives that train and improve frontier AI systems.
• Fluent or native proficiency in both English and Gujarati is essential.
• Possess strong judgment regarding language and content; capable of evaluating the accuracy, completeness, and appropriateness of AI responses.
• Ability to detect subtle errors, inconsistencies, and gaps in information.
• Consistently adhere to guidelines and quality standards.
• Capable of articulating reasoning effectively to both technical and non-technical audiences.
• Adaptable across various projects, task types, and client needs.
• Engagement as an independent contractor is required.
• Candidates must not be on an H1-B or STEM OPT visa.
• Nice-to-have: experience in adversarial machine learning, including knowledge of jailbreak datasets, prompt injections, RLHF/DPO attacks, and model extraction.
• Nice-to-have: background in cybersecurity, encompassing penetration testing, exploit development, and reverse engineering.
• Nice-to-have: experience in socio-technical risk assessment, including probing for harassment/disinformation, abuse analysis, and conversational AI testing.
• Nice-to-have: background in psychology, acting, or writing to foster unconventional adversarial thinking.
• Fully remote position.
• Flexible working hours; manage your own schedule.
• Receive weekly payments through Stripe or Wise.
• Optional participation in higher-sensitivity projects.
• Access to clear guidelines and wellness resources for working with sensitive content.
• Competitive compensation.
• Reasonable accommodations available upon request.
• Referral opportunity with earnings of up to $90 for each successful referral (subject to limits).
Mercor
RR Donnelley
MaintainX
Get handpicked remote jobs straight to your inbox weekly.