
AI Safety Expert, English, Assamese
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Engage in red-teaming of conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
• Generate reproducible reports, datasets, and attack scenarios for clients.
• Evaluate AI outputs related to sensitive subjects, including bias, misinformation, or harmful behaviors.
• Identify vulnerabilities that automated tests may overlook.
• Enhance evaluation coverage and minimize unexpected outcomes in production.
• Fortify customer AI systems through adversarial testing.
• Proficiency or native fluency in both English and Assamese.
• Strong discernment regarding language and content, particularly in assessing the accuracy, completeness, and appropriateness of AI responses.
• Capability to articulate reasoning clearly to both technical and non-technical audiences.
• Keen attention to detail to identify subtle errors, inconsistencies, and omissions.
• Consistent adherence to guidelines and quality standards.
• Flexibility to adapt across various projects, task types, and client requirements.
• Status as an independent contractor.
• Candidates must not be on an H-1B or STEM OPT visa.
• Preferred: experience in adversarial machine learning, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
• Preferred: background in cybersecurity, including penetration testing, exploit development, or reverse engineering.
• Preferred: experience with socio-technical risks, such as harassment/disinformation probing, abuse analysis, or conversational AI testing.
• Preferred: skills in psychology, acting, or writing to encourage unconventional adversarial thinking.
• Completely remote position that allows for flexible scheduling.
• Weekly payments processed through Stripe or Wise based on services provided.
• Project timelines may be adjusted, extended, or concluded early based on requirements and performance.
• Participation in higher-sensitivity projects is optional.
• Well-defined guidelines and wellness resources available for sensitive-content projects.
• Competitive compensation.
• Opportunity to collaborate with leading researchers.
• Referral program offering up to $90 for each successful referral, with no cap on the number of referrals (restrictions may apply).
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.