
AI Safety Expert β English, Bengali
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team evaluations of conversational AI models and agents using jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by annotating failures, identifying vulnerabilities, and highlighting systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Analyze AI outputs related to sensitive subjects such as bias, misinformation, or harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Enhance evaluation coverage and minimize unexpected issues in production.
β’ Assist Mercor customers in improving the safety, robustness, and reliability of AI systems.
β’ Fluent or native proficiency in both English and Bengali.
β’ Strong discernment regarding language and content.
β’ Capability to evaluate whether AI responses are accurate, complete, and suitable, and to articulate the reasoning behind assessments.
β’ Keen ability to detect subtle errors, inconsistencies, and omissions.
β’ Consistent adherence to guidelines and quality standards.
β’ Skill in clearly explaining reasoning to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects, task types, and client needs.
β’ Status as an independent contractor.
β’ Must not require H1-B or STEM OPT sponsorship.
β’ Preferred: experience in adversarial machine learning, including exposure to jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Preferred: background in cybersecurity, encompassing penetration testing, exploit development, or reverse engineering.
β’ Preferred: experience in socio-technical risks, such as harassment/disinformation probing, abuse analysis, or conversational AI evaluation.
β’ Preferred: qualifications in psychology, acting, or writing for innovative adversarial thinking.
β’ Fully remote position.
β’ Flexible schedule tailored to individual preferences.
β’ Weekly compensation via Stripe or Wise.
β’ Optional participation in higher-sensitivity projects.
β’ Clear guidelines and wellness resources available for sensitive-content tasks.
β’ Competitive salary.
β’ Opportunity to collaborate with top researchers.
β’ Referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.