
AI Safety Specialist, English, Dutch
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in New York.
• Execute adversarial assessments of conversational AI models and agents.
• Create jailbreaks, prompt-injection scenarios, misuse instances, and multi-turn manipulation techniques.
• Investigate models for vulnerabilities in various conversational and adversarial contexts.
• Utilize structured testing frameworks, recognized evaluation methodologies, taxonomies, benchmarks, and testing playbooks.
• Detect and categorize model failures and safety vulnerabilities associated with bias, misinformation, misuse, and detrimental model behaviors.
• Evaluate the severity, reproducibility, and practical implications of identified vulnerabilities.
• Generate high-quality human evaluation data from red teaming efforts.
• Annotate model failures and classify recognized vulnerabilities.
• Develop structured attack cases and accompanying evaluation materials.
• Clearly and reproducibly document adversarial scenarios.
• Generate reports, datasets, and structured findings for technical teams.
• Communicate identified risks to both technical and non-technical stakeholders.
• Maintain records of testing methodology, model behavior, and relevant failure patterns.
• Contribute to broader evaluation coverage across different models and use cases.
• Native-level fluency in both English and Dutch.
• Previous experience in AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing.
• Strong comprehension of conversational AI systems and model failure modes.
• Background in developing structured adversarial tests and evaluation frameworks.
• Ability to recognize subtle vulnerabilities and recurring behavioral patterns.
• Excellent analytical reasoning and written communication skills.
• Capability to document findings in a clear and reproducible manner.
• Comfort in working across evolving projects, scenarios, and evaluation frameworks.
• An educational background in computer science, cybersecurity, artificial intelligence, machine learning, linguistics, behavioral science, or a related field may be beneficial.
• Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation will also be considered.
• Practical red teaming experience is especially valuable.
• Experience in creating jailbreak or prompt-injection datasets.
• Familiarity with adversarial machine learning.
• Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation concepts.
• Cybersecurity experience related to penetration testing, exploit development, or reverse engineering.
• Background in analyzing abuse, harassment, misinformation, or other socio-technical risks.
• Experience in testing conversational AI systems.
• Strong creative writing, psychology, or behavioral-analysis skills applicable to adversarial testing.
• Previous experience in producing structured human data for AI evaluation.
• Native-level proficiency in English and Dutch is essential.
• Some projects may require reviewing text-based AI outputs related to sensitive topics such as bias, misinformation, or harmful behaviors.
• Work will not involve access to confidential or proprietary information from any employer, client, or institution.
• Flexible scheduling.
• Competitive hourly compensation ranging from $45 to $60.
• Weekly payments via Stripe or Wise.
• Participation in higher-sensitivity projects is optional, with relevant subject matter disclosed before involvement.
• Fully remote consulting opportunities.
• Project durations may be extended, shortened, or modified based on scope and performance.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.