
AI Safety Experts – English, Bengali
Posted Sep 4

Posted Sep 4
This is a fully remote position, open to applicants in United States.
• Conduct red-team evaluations of conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure testing remains consistent.
• Generate reproducible reports, datasets, and attack scenarios for clients.
• Analyze AI outputs related to sensitive issues such as bias, misinformation, or harmful conduct.
• Identify vulnerabilities that automated tests may overlook.
• Broaden evaluation coverage while minimizing unexpected production issues.
• Enhance customer AI systems through adversarial testing.
• Must possess native proficiency in English and Bengali.
• Previous experience in red teaming for AI adversarial work, cybersecurity, or socio-technical probing is essential.
• Capability to probe systems adversarially and challenge them to their limits.
• Familiarity with frameworks, taxonomies, benchmarks, or playbooks is required.
• Ability to clearly communicate risks to both technical and non-technical stakeholders.
• Must demonstrate adaptability across various projects and clients.
• Independent contractor status is required.
• Candidates must not hold H1-B or STEM OPT status.
• Preferred qualifications include expertise in adversarial ML, cybersecurity, socio-technical risk, and creative probing.
• Experience in adversarial ML involving jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction is advantageous.
• Cybersecurity background in penetration testing, exploit development, or reverse engineering is a plus.
• Experience in socio-technical areas involving harassment/disinformation probing, abuse analysis, or conversational AI testing is beneficial.
• Background in creative probing through psychology, acting, or writing is an asset.
• Fully remote position.
• Flexible work schedule that allows you to manage your own time.
• Weekly payments through Stripe or Wise.
• Project durations may vary based on needs and performance.
• Participation in higher-sensitivity projects is optional.
• Comprehensive guidelines and wellness resources for handling sensitive content.
• Competitive compensation.
• Opportunity to collaborate with leading researchers.
• Referral bonus of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.