
AI Safety Expert, English – Tamil
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in United States.
• Engage in red-teaming of conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
• Create human data by annotating failures, classifying vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing.
• Generate reproducible reports, datasets, and attack scenarios for clients.
• Analyze AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
• Discover vulnerabilities that automated tests may overlook.
• Broaden evaluation coverage and minimize unexpected issues in production.
• Foster customer confidence in AI safety by conducting adversarial probes of systems.
• Required fluency or native proficiency in both English and Tamil.
• Strong discernment regarding language and content; capable of evaluating the accuracy, completeness, and appropriateness of AI responses.
• Ability to detect subtle errors, inconsistencies, and omissions.
• Consistent adherence to taxonomies, benchmarks, playbooks, guidelines, and quality standards.
• Capable of articulating reasoning clearly to both technical and non-technical audiences.
• Flexibility to adapt across various projects, task types, and client needs.
• Engagement as an independent contractor is essential.
• Candidates must not be on H1-B or STEM OPT visas.
• Preferred qualifications: experience in adversarial machine learning, including work with jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
• Preferred qualifications: experience in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
• Preferred qualifications: experience in socio-technical risk, including probing for harassment/disinformation, abuse analysis, or testing conversational AI.
• Preferred qualifications: background in psychology, acting, or writing to facilitate unconventional adversarial thinking.
• Fully remote position.
• Flexible working hours; manage your own schedule.
• Weekly payments processed through Stripe or Wise.
• Access to wellness resources and clear guidelines for high-sensitivity projects.
• Gain professional experience in human data-driven AI red teaming.
• Play a direct role in enhancing the robustness, safety, and trustworthiness of AI systems.
• Competitive compensation.
• Opportunity to collaborate with leading researchers.
• Referral bonus of up to $90 for each successful referral, subject to limits.
Epoch AI
Wagmo
Mercor
Blend360
Get handpicked remote jobs straight to your inbox weekly.