
AI Safety Expert β English, Bengali
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Generate human data by annotating errors, classifying vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing protocols.
β’ Create reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs concerning sensitive subjects like bias, misinformation, and harmful behaviors.
β’ Assist in broadening evaluation coverage and enhancing client AI systems.
β’ Proficient/native proficiency in English and Bengali.
β’ Strong judgment regarding language and content; capable of assessing the accuracy, completeness, and appropriateness of AI responses and articulating the reasoning behind evaluations.
β’ Keen ability to detect subtle errors, inconsistencies, and omissions.
β’ Consistently adhere to guidelines and quality standards.
β’ Ability to communicate reasoning effectively to both technical and non-technical audiences.
β’ Flexibility across various projects, task types, and clients.
β’ Must operate as an independent contractor.
β’ Must not require H1-B or STEM OPT sponsorship.
β’ Preferred: experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Preferred: experience in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
β’ Preferred: experience in socio-technical risk analysis, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
β’ Preferred: background in psychology, acting, or writing for unconventional adversarial thought processes.
β’ Fully remote position.
β’ Flexible working hours to manage your schedule.
β’ Weekly payments through Stripe or Wise.
β’ Optional participation in higher-sensitivity projects.
β’ Clear guidelines and wellness resources available for sensitive-content tasks.
β’ Competitive compensation.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Collaborate with leading experts in the field.
β’ Referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.