AI Safety Expert – English, Bengali

atMercorRemoteUS flagUnited StatesFreelanceArtificial IntelligenceMid-levelSenior$20 – $22/hour

Posted Aug 25

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Conduct red-team assessments on conversational AI models and agents by employing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.

β€’ Create human data through the annotation of failures, classification of vulnerabilities, and identification of systemic risks.

β€’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing protocols.

β€’ Generate reproducible reports, datasets, and attack case scenarios for clients.

β€’ Analyze AI outputs related to sensitive subjects such as bias, misinformation, or harmful behaviors.

β€’ Identify vulnerabilities that automated testing may overlook.

β€’ Broaden evaluation coverage and minimize unexpected issues in production.

β€’ Enhance customer AI systems through adversarial testing methodologies.


⛳️ Requirements

β€’ Native or fluent proficiency in both English and Bengali is mandatory.

β€’ Previous experience in red teaming, particularly in AI adversarial contexts, cybersecurity, or socio-technical investigations.

β€’ Capability to probe AI systems adversarially and test their limits.

β€’ Familiarity with frameworks, taxonomies, benchmarks, or playbooks for methodical testing.

β€’ Skill in articulating risks clearly to both technical and non-technical audiences.

β€’ Flexibility to adapt across various projects and clients.

β€’ Preferred expertise includes adversarial machine learning, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation probing, abuse analysis, conversational AI testing, psychology, acting, or writing.

β€’ Must hold independent contractor status.

β€’ Candidates on H1-B or STEM OPT visas will not be considered.


🏝️ Benefits

β€’ Enjoy the flexibility of fully remote work on your own schedule.

β€’ Receive weekly payments through Stripe or Wise.

β€’ Optional participation in projects with higher sensitivity levels.

β€’ Access to clear content guidelines prior to engaging with sensitive subjects.

β€’ Benefit from wellness resources.

β€’ Competitive compensation package.

β€’ Earn referral payments of up to $90 for each successful referral.

β€’ Collaborate with leading researchers in the field.

β€’ Reasonable accommodations available upon request.

People also viewed

Sigma AI21 hours ago

Marathi Linguistic Projects

IN flagIndia OnlyFreelanceArtificial Intelligence
ApplyView job
Curri22 hours ago

Director, Data – AI

US flagCalifornia OnlyFull-timeArtificial Intelligence$220k – $260k/year
ApplyView job
Mercor22 hours ago

AI Safety Expert, English, Marathi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor22 hours ago

AI Safety Expert, English – Assamese

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Staples Promotional Products22 hours ago

Manager, Sales Intelligence – AI Strategy

US flagMassachusetts OnlyFull-timeArtificial Intelligence
ApplyView job
Language Inspired22 hours ago

Language Quality, AI Specialist

EuropeFull-timeArtificial Intelligence
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers