
AI Safety Expert, English, Dutch
Posted 7 hours ago

Posted 7 hours ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments on conversational AI models and agents by utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create human data through the annotation of failures, classification of vulnerabilities, and identification of systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing procedures.
• Generate reproducible reports, datasets, and attack case studies for clients.
• Analyze AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
• Detect vulnerabilities that automated testing may overlook.
• Broaden evaluation coverage to minimize unexpected issues during production.
• Enhance customer AI systems through adversarial testing methodologies.
• Proficient/native proficiency in both English and Dutch.
• Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical investigations.
• Capability to challenge AI systems adversarially and push them to their limits.
• Familiarity with frameworks, taxonomies, benchmarks, or playbooks for structured testing approaches.
• Ability to articulate risks clearly to both technical and non-technical audiences.
• Flexibility to adapt across various projects and clientele.
• Status as an independent contractor.
• Must not require H1-B or STEM OPT sponsorship.
• Preferred specialties: adversarial machine learning, cybersecurity, socio-technical risk, or innovative probing techniques.
• Adversarial ML experience may involve working with jailbreak datasets, prompt injections, reinforcement learning from human feedback (RLHF)/differentiable programming optimization (DPO) attacks, or model extraction.
• Cybersecurity experience may include penetration testing, exploit development, or reverse engineering.
• Socio-technical risk experience may encompass harassment/disinformation probing, abuse analysis, or testing of conversational AI systems.
• Creative probing experience may involve psychology, acting, or writing skills.
• Fully remote position.
• Flexible working hours; tasks can be completed according to your own schedule.
• Weekly payments processed through Stripe or Wise.
• Project durations may be extended, shortened, or concluded early based on requirements and performance.
• Access to wellness resources and clear guidelines for projects with higher sensitivity.
• Opportunity to gain experience in human data-driven AI red teaming.
• Play a direct role in enhancing the robustness, safety, and trustworthiness of AI systems.
• Competitive compensation offered.
• Referral bonuses of up to $250 available for each successful referral, subject to certain limits.
• Reasonable accommodations can be requested as needed.
BPCS, Comprehensive marketing solutions, ltd.
Tech Mahindra
PwC
Get handpicked remote jobs straight to your inbox weekly.