AI Safety Experts, English – Gujarati

atMercorRemoteUS flagUnited StatesFreelanceArtificial IntelligenceMid-levelSenior$16 – $22/hour

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Conduct red-team evaluations on conversational AI models and agents through jailbreak techniques, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.

• Generate human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.

• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.

• Create reproducible reports, datasets, and attack cases for client use.

• Analyze AI outputs related to sensitive subjects such as bias, misinformation, and harmful behavior.

• Expose vulnerabilities that automated tests may overlook.

• Enhance evaluation coverage and minimize unexpected issues during production.

• Collaborate on initiatives that train and improve frontier AI systems.


⛳️ Requirements

• Fluent or native proficiency in both English and Gujarati is essential.

• Possess strong judgment regarding language and content; capable of evaluating the accuracy, completeness, and appropriateness of AI responses.

• Ability to detect subtle errors, inconsistencies, and gaps in information.

• Consistently adhere to guidelines and quality standards.

• Capable of articulating reasoning effectively to both technical and non-technical audiences.

• Adaptable across various projects, task types, and client needs.

• Engagement as an independent contractor is required.

• Candidates must not be on an H1-B or STEM OPT visa.

• Nice-to-have: experience in adversarial machine learning, including knowledge of jailbreak datasets, prompt injections, RLHF/DPO attacks, and model extraction.

• Nice-to-have: background in cybersecurity, encompassing penetration testing, exploit development, and reverse engineering.

• Nice-to-have: experience in socio-technical risk assessment, including probing for harassment/disinformation, abuse analysis, and conversational AI testing.

• Nice-to-have: background in psychology, acting, or writing to foster unconventional adversarial thinking.


🏝️ Benefits

• Fully remote position.

• Flexible working hours; manage your own schedule.

• Receive weekly payments through Stripe or Wise.

• Optional participation in higher-sensitivity projects.

• Access to clear guidelines and wellness resources for working with sensitive content.

• Competitive compensation.

• Reasonable accommodations available upon request.

• Referral opportunity with earnings of up to $90 for each successful referral (subject to limits).

People also viewed

Cresta19 hours ago

AI Strategist

US flagUnited States OnlyFull-timeArtificial Intelligence
ApplyView job
Mercor19 hours ago

AI Safety Expert – English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
RR Donnelley19 hours ago

AI Workflow Engineer

US flagIllinois OnlyFreelanceArtificial Intelligence$107k – $171.2k/year
ApplyView job
MaintainX19 hours ago

Senior Director, MaintainX AI

US flagCalifornia OnlyFull-timeArtificial Intelligence
ApplyView job
Motorola Solutions21 hours ago

AI Manager – Language Models

US flagArizona, +9 more statesFull-timeArtificial Intelligence$240k – $265k/year
ApplyView job
Beglaubigt.de (YC F24)22 hours ago

Operations & AI Analyst Intern

DE flagGermany OnlyInternshipArtificial Intelligence
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers