AI Safety Expert – English, Tamil

atMercorRemoteUS flagUnited StatesFreelanceArtificial IntelligenceMid-levelSenior$16 – $22/hour

Posted 22 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Conduct red-team assessments on conversational AI models and agents through jailbreak techniques, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.

• Generate human-centric data by annotating failures, classifying vulnerabilities, and identifying systemic risks.

• Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing practices.

• Create reproducible reports, datasets, and attack scenarios for clients.

• Evaluate AI outputs related to sensitive subjects like bias, misinformation, and harmful behaviors.

• Identify vulnerabilities that automated testing may overlook.

• Enhance evaluation coverage and minimize unexpected issues during production.

• Fortify customer AI systems through adversarial testing.


⛳️ Requirements

• Proficiency in both English and Tamil is essential.

• Strong discernment regarding language and content.

• Capability to evaluate AI responses for accuracy, completeness, and appropriateness, with the ability to articulate reasoning.

• Skill in identifying subtle errors, inconsistencies, and gaps in information.

• Consistent adherence to guidelines and quality standards.

• Ability to communicate reasoning effectively to both technical and non-technical audiences.

• Flexibility to adapt across various projects, task types, and client needs.

• Status as an independent contractor is required.

• H1-B and STEM OPT candidates will not be accommodated.

• Preferred: experience in adversarial machine learning, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.

• Preferred: background in cybersecurity, encompassing penetration testing, exploit development, or reverse engineering.

• Preferred: experience in socio-technical risk, including harassment/disinformation probing, abuse analysis, or conversational AI testing.

• Preferred: creative probing experience in fields such as psychology, acting, or writing.


🏝️ Benefits

• Fully remote position.

• Flexible working hours, allowing you to manage your own schedule.

• Receive weekly payments via Stripe or Wise.

• Optional participation in high-sensitivity projects.

• Access to clear guidelines and wellness resources for working with sensitive content.

• Competitive compensation.

• Opportunity to collaborate with leading researchers in the field.

• Referral bonuses of up to $90 for each successful referral, subject to certain limits.

• Reasonable accommodations available upon request.

People also viewed

Mercor13 hours ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor13 hours ago

AI Safety Expert – English, Telugu

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
snipKI20 hours ago

Community Manager – AI Community

DE flagGermany OnlyPart-timeArtificial Intelligence
ApplyView job
Mercor22 hours ago

AI Safety Expert – English, Finnish

US flagUnited States OnlyFreelanceArtificial Intelligence$48 – $62/hour
ApplyView job
Mercor22 hours ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor1 day ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers