
AI Safety Expert β English, Finnish
Posted Sep 14

Posted Sep 14
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Generate human data by annotating failures, classifying vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to maintain consistency in testing.
β’ Create reproducible reports, datasets, and attack cases tailored for clients.
β’ Investigate AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Discover vulnerabilities that automated testing may overlook.
β’ Broaden evaluation coverage and minimize unexpected issues in production.
β’ Assist clients in enhancing the safety and robustness of their AI systems.
β’ Must possess fluent/native proficiency in both English and Finnish.
β’ Previous experience in red teaming related to AI adversarial tasks, cybersecurity, or socio-technical probing is essential.
β’ Capability to challenge AI systems adversarially and test their limits.
β’ Familiarity with frameworks, taxonomies, benchmarks, or playbooks for structured testing is required.
β’ Proficient in articulating risks to both technical and non-technical stakeholders.
β’ Demonstrated adaptability across various projects and clientele.
β’ Available as an independent contractor.
β’ Must be able to work without access to confidential or proprietary information from other employers, clients, or institutions.
β’ H1-B and STEM OPT candidates are not eligible.
β’ Preferred specialties include adversarial machine learning, cybersecurity, socio-technical risk, or innovative probing techniques.
β’ Fully remote position.
β’ Flexible, self-directed work schedule.
β’ Weekly payments processed through Stripe or Wise based on services provided.
β’ Project durations may be adjusted based on needs and performance.
β’ Participation in higher-sensitivity projects is voluntary.
β’ Clear content guidelines and wellness resources available.
β’ Reasonable accommodations can be requested.
β’ Up to $250 referral bonus for each successful referral, with no limit on the number of referrals (certain restrictions may apply).
β’ Collaborate with leading researchers in the field.
β’ Gain experience in human data-driven AI red teaming.
β’ Contribute to the development of safer, more robust, and trustworthy AI systems.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.