
AI Safety Specialist, English, Finnish
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in New York.
• Perform adversarial testing on conversational AI models and agents.
• Design jailbreaks, prompt-injection scenarios, misuse cases, and multi-turn manipulation tactics.
• Investigate models for vulnerabilities in various conversational and adversarial situations.
• Implement systematic testing frameworks and recognized evaluation methodologies.
• Detect and categorize model failures and safety vulnerabilities.
• Assess bias, misinformation, misuse, and potentially harmful behaviors of models.
• Identify recurring or systemic patterns within model responses.
• Evaluate the severity, reproducibility, and practical implications of vulnerabilities.
• Generate high-quality human evaluation data from red teaming exercises.
• Annotate model failures and classify vulnerabilities.
• Develop structured attack cases and accompanying evaluation materials.
• Ensure consistency across repeated assessments and datasets.
• Clearly and reproducibly document adversarial scenarios.
• Create reports, datasets, and structured findings for technical teams.
• Communicate identified risks to both technical and non-technical stakeholders.
• Record testing methodologies, model behaviors, and pertinent failure patterns.
• Contribute to enhanced evaluation coverage across models and use cases.
• Native-level proficiency in both English and Finnish is mandatory.
• Previous experience in AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing.
• Strong knowledge of conversational AI systems and their failure modes.
• Experience in creating structured adversarial tests and evaluation frameworks.
• Ability to detect subtle vulnerabilities and recurring behavioral patterns.
• Excellent analytical reasoning and written communication skills.
• Capacity to document findings clearly and reproducibly.
• Comfort in navigating varying projects, scenarios, and evaluation frameworks.
• An educational background in computer science, cybersecurity, artificial intelligence, machine learning, linguistics, behavioral science, or a related field may be advantageous.
• Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation may be accepted.
• Practical experience in red teaming is especially beneficial.
• Familiarity with adversarial machine learning.
• Understanding of RLHF, DPO, model extraction, or related AI training and evaluation concepts.
• Cybersecurity experience related to penetration testing, exploit development, or reverse engineering.
• Experience in testing conversational AI systems.
• Previous experience in generating structured human data for AI evaluation.
• Fully remote work environment.
• Flexible scheduling options.
• Competitive hourly pay ranging from $45 to $60 per hour.
• Weekly payments processed via Stripe or Wise.
• Participation in high-sensitivity projects is optional.
• Project duration may be extended, shortened, or adjusted based on scope and performance.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.