
AI Safety Expert – English, Assamese
Posted Aug 13

Posted Aug 13
This is a fully remote position, open to applicants in United States.
• Conduct red team assessments on conversational AI models and agents by utilizing jailbreak techniques, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create high-quality human data through annotating failures, identifying vulnerabilities, and highlighting systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing methodologies.
• Generate reproducible reports, datasets, and attack cases tailored for clients.
• Examine AI outputs related to sensitive subjects including bias, misinformation, and detrimental behaviors.
• Identify vulnerabilities that automated testing may overlook.
• Broaden evaluation scope and minimize unexpected issues in production.
• Enhance customer AI systems to make them more robust, secure, and reliable.
• Collaborate with top AI researchers on initiatives focused on training and improving cutting-edge AI systems.
• Proficiency in both English and Assamese is essential.
• Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical analysis.
• Capability to probe systems adversarially and challenge them to their limits.
• Familiarity with structured frameworks, taxonomies, benchmarks, or playbooks.
• Ability to effectively communicate risks to both technical and non-technical stakeholders.
• Flexibility to adapt across various projects and client needs.
• Expertise or interest in areas such as jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation probing, abuse analysis, conversational AI testing, psychology, acting, or writing.
• Must operate as an independent contractor.
• H1-B and STEM OPT candidates are not eligible.
• Work remotely from anywhere.
• Enjoy a flexible work schedule and the freedom to manage your own time.
• Receive weekly payments through Stripe or Wise.
• Competitive compensation.
• Collaborate with leading researchers in the field.
• Gain experience in human data-driven AI red teaming.
• Access to wellness resources and clear guidelines for projects requiring higher sensitivity.
• Participation in higher-sensitivity projects is optional.
• Reasonable accommodations available upon request.
• Earn referral payments of up to $90 for each successful referral.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.