
AI Safety Expert, English – Kannada
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team evaluations of conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Generate human data by documenting failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing.
• Create reproducible reports, datasets, and attack scenarios that clients can act upon.
• Review AI outputs related to sensitive subjects, including bias, misinformation, and harmful behaviors.
• Identify vulnerabilities overlooked by automated testing.
• Broaden evaluation coverage and minimize unexpected production issues.
• Collaborate on projects to train and enhance AI systems for Mercor’s AI lab and its enterprise clients.
• Native or fluent proficiency in both English and Kannada is essential.
• Strong judgment regarding language and content, particularly in assessing the accuracy, completeness, and appropriateness of AI responses.
• Ability to articulate reasoning clearly to both technical and non-technical audiences.
• Keen attention to subtle errors, inconsistencies, and gaps.
• Capability to consistently adhere to taxonomies, benchmarks, playbooks, guidelines, and quality standards.
• Flexibility to adapt across different projects, task types, and clients.
• Must hold independent contractor status.
• Candidates must not be on an H1-B or STEM OPT visa.
• Preferred: experience in adversarial machine learning, including work with jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
• Preferred: background in cybersecurity, encompassing penetration testing, exploit development, or reverse engineering.
• Preferred: experience in socio-technical risk, including harassment/disinformation probing, abuse analysis, or testing conversational AI.
• Preferred: skills in psychology, acting, or writing that encourage unconventional adversarial thinking.
• Fully remote position that allows for flexible scheduling.
• Weekly payments through Stripe or Wise based on services provided.
• Opportunity to gain experience in human data-driven AI red teaming at the forefront of safety.
• Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
• Access to wellness resources and clear guidelines for projects involving higher sensitivity.
• Reasonable accommodations available upon request.
• Referral bonuses of up to $90 for each successful referral, subject to specific restrictions.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.