AI Safety Experts – English, Marathi

atMercorRemoteUS flagUnited StatesFreelanceArtificial IntelligenceMid-levelSenior$16 – $22/hour

Posted 4 days ago

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Conduct red-team assessments on conversational AI models and agents through techniques such as jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.

β€’ Create human data by annotating failures, classifying vulnerabilities, and identifying systemic risks.

β€’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing procedures.

β€’ Generate reproducible reports, datasets, and attack scenarios for clients.

β€’ Evaluate AI outputs related to sensitive subjects like bias, misinformation, and harmful behaviors.

β€’ Identify vulnerabilities that automated testing may overlook.

β€’ Enhance evaluation coverage and minimize unexpected issues in production.

β€’ Fortify customer AI systems through adversarial testing.


⛳️ Requirements

β€’ Native fluency in both English and Marathi is essential.

β€’ Strong discernment regarding language and content is required.

β€’ Capability to assess the accuracy, completeness, and appropriateness of AI responses and articulate the reasoning behind evaluations.

β€’ Meticulous attention to detail, ensuring subtle errors, inconsistencies, and gaps are identified.

β€’ Consistent adherence to guidelines and quality standards is necessary.

β€’ Ability to convey reasoning clearly to both technical and non-technical audiences.

β€’ Flexibility to adapt across various projects, task types, and client needs.

β€’ Must hold independent contractor status.

β€’ H1-B and STEM OPT candidates are not eligible.

β€’ Preferred: experience in adversarial ML, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.

β€’ Preferred: background in cybersecurity, including penetration testing, exploit development, or reverse engineering.

β€’ Preferred: experience in socio-technical risk, such as harassment/disinformation probing, abuse analysis, or testing of conversational AI.

β€’ Preferred: skills in psychology, acting, or writing for innovative adversarial thinking.


🏝️ Benefits

β€’ Fully remote position.

β€’ Flexible work schedule; complete tasks at your convenience.

β€’ Weekly payments through Stripe or Wise.

β€’ Project durations may vary based on needs and performance.

β€’ Optional participation in higher-sensitivity projects.

β€’ Clear guidelines and wellness resources available for sensitive-content tasks.

β€’ Reasonable accommodations provided upon request.

β€’ Competitive compensation.

β€’ Opportunity to collaborate with leading researchers.

β€’ Referral bonuses of up to $90 for each successful referral, subject to limitations.

People also viewed

WON.ai16 hours ago

AI Strategist

AR flagArgentina OnlyFreelanceArtificial Intelligence
ApplyView job
The College Board17 hours ago

Director, AI Assisted Solutions

US flagUnited States OnlyFull-timeArtificial Intelligence$88k – $135k/year
ApplyView job
Mercor18 hours ago

AI Safety Experts – English, Assamese

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor18 hours ago

AI Safety Experts – English, Marathi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor18 hours ago

AI Safety Expert – English, Bengali

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Gartner18 hours ago

Director, Analyst – AI Technology Economics

GB flagUnited Kingdom OnlyFull-timeArtificial Intelligence
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers