
AI Safety Expert – English, Bengali
Posted 12 hours ago

Posted 12 hours ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments on conversational AI models and agents through jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Generate human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to maintain testing consistency.
• Create reproducible reports, datasets, and attack scenarios for clients.
• Identify vulnerabilities that automated testing may overlook.
• Broaden evaluation coverage and minimize unexpected issues in production.
• Assist Mercor’s AI-lab and corporate clients in enhancing the safety, robustness, and trustworthiness of frontier AI systems.
• Native proficiency in both English and Bengali.
• Previous experience in red teaming related to AI adversarial work, cybersecurity, or socio-technical probing.
• Capability to probe AI systems in an adversarial manner and push them to their limits.
• Familiarity with using frameworks, taxonomies, benchmarks, or structured playbooks.
• Ability to articulate risks clearly to both technical and non-technical audiences.
• Flexibility to adapt across different projects and client needs.
• Preferred: experience in adversarial ML, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
• Preferred: cybersecurity experience, such as penetration testing, exploit development, or reverse engineering.
• Preferred: experience in socio-technical risk assessment, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
• Preferred: background in psychology, acting, or writing to encourage unconventional adversarial thinking.
• Must meet eligibility criteria; H-1B and STEM OPT candidates are not eligible.
• Engagement as an independent contractor.
• Fully remote position.
• Flexible working hours / ability to work according to your own schedule.
• Weekly payments via Stripe or Wise.
• Project durations may be extended, reduced, or concluded prematurely based on needs and performance.
• Participation in higher-sensitivity projects is optional.
• Clear content guidelines and wellness resources available for sensitive-topic work.
• Reasonable accommodations provided upon request.
• Competitive compensation.
• Referral bonuses of up to $90 for each successful referral.
Mercor
Riva Scientific
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.