
AI Safety Expert, English – Dutch
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Conduct red team assessments on conversational AI models and agents through techniques such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create high-quality human data by annotating failures, identifying vulnerabilities, and flagging systemic risks.
• Utilize taxonomies, benchmarks, and playbooks to ensure consistency in testing processes.
• Generate reproducible reports, datasets, and attack cases for clients.
• Identify vulnerabilities that automated testing may overlook.
• Broaden evaluation coverage and minimize unexpected issues in production.
• Enhance the safety, robustness, and trustworthiness of customer AI systems.
• Required native fluency in both English and Dutch.
• Previous experience in red teaming related to AI adversarial operations, cybersecurity, or socio-technical probing.
• Capability to probe systems in an adversarial manner and challenge them to their limits.
• Familiarity with frameworks or benchmarks for structured testing methodologies.
• Proficient in clearly communicating risks to both technical and non-technical audiences.
• Flexibility to adapt across various projects and clients.
• Nice-to-have: experience in adversarial machine learning, including work with jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
• Nice-to-have: background in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
• Nice-to-have: experience in socio-technical risk assessment, including probing for harassment/disinformation, abuse analysis, or conversational AI testing.
• Nice-to-have: experience in psychology, acting, or writing to foster unconventional adversarial thinking.
• Must be capable of working as an independent contractor.
• H1-B and STEM OPT candidates are not eligible.
• Fully remote position.
• Flexible scheduling options.
• Participation in higher-sensitivity projects is optional.
• Clear guidelines and wellness resources available for managing sensitive content.
• Weekly payment options through Stripe or Wise.
• Competitive compensation.
• Opportunity to work alongside leading researchers in the field.
• Referral bonus of up to $250 for successful referrals.
Progressive Leasing
apna
apna
Texas Research International
Get handpicked remote jobs straight to your inbox weekly.