
AI Safety Expert – English, Finnish
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in United States.
• Conduct red team assessments on conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Document failures, categorize vulnerabilities, and highlight systemic risks.
• Adhere to established taxonomies, benchmarks, and playbooks to ensure consistency in testing.
• Create reproducible reports, datasets, and attack scenarios for clients.
• Identify vulnerabilities that automated testing may overlook.
• Broaden evaluation coverage to minimize unexpected issues during production.
• Collaborate with top researchers and engage in projects aimed at training and refining AI systems.
• Proficient in both English and Finnish, with fluency or native-level skills required.
• Previous experience in red teaming related to AI adversarial tasks, cybersecurity, or socio-technical probing.
• Capability to interrogate AI systems adversarially and push them to their limits.
• Experience utilizing frameworks or benchmarks for systematic testing.
• Ability to communicate risks effectively to both technical and non-technical audiences.
• Flexibility to adapt across various projects and clientele.
• Engagement as an independent contractor.
• Not eligible if you are an H1-B or STEM OPT candidate.
• Nice-to-have: experience in adversarial machine learning, including familiarity with jailbreak datasets, prompt injections, RLHF/DPO attacks, or model extraction.
• Nice-to-have: cybersecurity background, including penetration testing, exploit development, or reverse engineering.
• Nice-to-have: experience in assessing socio-technical risks, such as harassment/disinformation probing, abuse analysis, or conversational AI testing.
• Nice-to-have: background in psychology, acting, or writing to foster unconventional adversarial thinking.
• Completely remote position.
• Flexible scheduling according to your preferences.
• Weekly compensation processed through Stripe or Wise.
• Access to wellness resources and clear protocols for sensitive projects.
• Opportunity to gain experience in human data-driven AI red teaming.
• Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
• Competitive salary.
• Referral bonus of up to $250 for each successful referral.
Mercor
Mercor
snipKI
Mercor
Get handpicked remote jobs straight to your inbox weekly.