
AI Safety Experts – English, Marathi
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments on conversational AI models and agents utilizing jailbreak techniques, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
• Create human-generated data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing methodologies.
• Develop reproducible reports, datasets, and attack scenarios for clients.
• Review AI outputs related to sensitive issues such as bias, misinformation, and harmful behaviors.
• Identify vulnerabilities that automated testing may overlook.
• Broaden evaluation coverage and minimize unexpected outcomes in production.
• Assist Mercor clients in enhancing the safety, robustness, and trustworthiness of their AI systems.
• Must possess native fluency in both English and Marathi.
• Strong judgment regarding language use and content; capable of evaluating the accuracy, completeness, and appropriateness of AI responses.
• Proficient in spotting subtle errors, inconsistencies, and gaps in information.
• Consistently able to adhere to guidelines and quality standards.
• Capable of articulating reasoning effectively to both technical and non-technical audiences.
• Flexible in adapting to various projects, task types, and client needs.
• Must hold independent contractor status.
• Should not require H1-B or STEM OPT support.
• Nice-to-have: experience in adversarial machine learning, including working with jailbreak datasets, prompt injections, RLHF/DPO attacks, or model extraction.
• Nice-to-have: background in cybersecurity, including penetration testing, exploit development, or reverse engineering.
• Nice-to-have: experience in socio-technical risks, such as harassment/disinformation probing, abuse analysis, or testing of conversational AI.
• Nice-to-have: experience in psychology, acting, or writing for unconventional adversarial thinking.
• Fully remote position that allows for flexible scheduling.
• Weekly compensation through Stripe or Wise.
• Optional participation in higher-sensitivity projects.
• Clear guidelines and wellness resources available for handling sensitive content.
• Competitive remuneration.
• Opportunity to collaborate with leading researchers in the field.
• Referral bonuses of up to $90 for each successful referral (conditions apply).
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.