
AI Safety Expert β English, Bengali
Posted Aug 25

Posted Aug 25
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents by employing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data through the annotation of failures, classification of vulnerabilities, and identification of systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing protocols.
β’ Generate reproducible reports, datasets, and attack case scenarios for clients.
β’ Analyze AI outputs related to sensitive subjects such as bias, misinformation, or harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Broaden evaluation coverage and minimize unexpected issues in production.
β’ Enhance customer AI systems through adversarial testing methodologies.
β’ Native or fluent proficiency in both English and Bengali is mandatory.
β’ Previous experience in red teaming, particularly in AI adversarial contexts, cybersecurity, or socio-technical investigations.
β’ Capability to probe AI systems adversarially and test their limits.
β’ Familiarity with frameworks, taxonomies, benchmarks, or playbooks for methodical testing.
β’ Skill in articulating risks clearly to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects and clients.
β’ Preferred expertise includes adversarial machine learning, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation probing, abuse analysis, conversational AI testing, psychology, acting, or writing.
β’ Must hold independent contractor status.
β’ Candidates on H1-B or STEM OPT visas will not be considered.
β’ Enjoy the flexibility of fully remote work on your own schedule.
β’ Receive weekly payments through Stripe or Wise.
β’ Optional participation in projects with higher sensitivity levels.
β’ Access to clear content guidelines prior to engaging with sensitive subjects.
β’ Benefit from wellness resources.
β’ Competitive compensation package.
β’ Earn referral payments of up to $90 for each successful referral.
β’ Collaborate with leading researchers in the field.
β’ Reasonable accommodations available upon request.
Sigma AI
Curri
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.