
AI Safety Expert, English, Thai
Posted 20 hours ago

Posted 20 hours ago
This is a fully remote position, open to applicants in United States.
• Engage in red teaming for conversational AI models and agents
• Test for jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations
• Create high-quality human data through annotating failures, classifying vulnerabilities, and identifying systemic risks
• Utilize taxonomies, benchmarks, and playbooks to maintain consistency in testing
• Generate reproducible reports, datasets, and attack cases for clients
• Identify vulnerabilities that automated testing may overlook
• Broaden evaluation coverage and minimize unexpected production issues
• Analyze AI outputs related to sensitive subjects such as bias, misinformation, or harmful behaviors
• Collaborate with leading researchers and assist in the training and enhancement of cutting-edge AI systems
• Proficient/native fluency in both English and Thai
• Previous experience in red teaming, specifically in AI adversarial work, cybersecurity, or socio-technical exploration
• Capability to challenge AI systems using adversarial inputs
• Familiarity with jailbreaks, prompt injections, misuse scenarios, bias exploitation, or multi-turn manipulations
• Skill in annotating failures, classifying vulnerabilities, and identifying systemic risks
• Competence in adhering to taxonomies, benchmarks, and playbooks
• Ability to produce reproducible reports, datasets, and attack cases
• Proficiency in clearly explaining risks to both technical and non-technical audiences
• Flexibility to adapt across various projects and clients
• Preferred specialties include adversarial ML, cybersecurity, socio-technical risks, or creative probing
• Must be capable of working as an independent contractor
• H1-B and STEM OPT candidates are not eligible
• Fully remote position
• Flexible working hours / ability to set your own schedule
• Weekly compensation through Stripe or Wise
• Project durations may be adjusted based on needs and performance outcomes
• Access to wellness resources and clear protocols for sensitive projects
• Direct involvement in human data-driven AI red teaming
• Chance to contribute to the development of safer, more robust, and trustworthy AI systems
• Competitive compensation
• Referral bonus of up to $140 for each successful referral
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.