
AI Safety Expert, English, Finnish
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
β’ Engage in red team activities with conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Adhere to taxonomies, benchmarks, and playbooks to ensure consistent testing procedures.
β’ Generate reproducible reports, datasets, and attack cases for clients.
β’ Identify vulnerabilities that automated tests may overlook.
β’ Deliver reproducible artifacts that enhance the robustness of customer AI systems.
β’ Broaden evaluation coverage by testing additional scenarios and minimizing production surprises.
β’ Proficient/native fluency in both English and Finnish.
β’ Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical probing.
β’ Capability to probe AI systems adversarially, including the use of jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
β’ Skill in generating human data through failure annotation, vulnerability classification, and systemic risk identification.
β’ Ability to adhere to taxonomies, benchmarks, and playbooks.
β’ Expertise in producing reproducible reports, datasets, and attack cases.
β’ Proficiency in clearly communicating risks to both technical and non-technical stakeholders.
β’ Flexibility to adapt across various projects and customer requirements.
β’ Involvement in higher-sensitivity projects is optional.
β’ Clear guidelines established for handling sensitive-content work.
β’ Access to wellness resources.
Genesys
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.