
AI Safety Specialist, English, Swedish
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in New York.
• Perform adversarial assessments on conversational AI models and agents.
• Design jailbreaks, prompt-injection scenarios, misuse cases, and multi-turn manipulation techniques.
• Investigate models for weaknesses not detected by automated evaluation systems.
• Analyze model behavior across various conversational and adversarial situations.
• Utilize systematic testing frameworks and recognized evaluation methodologies.
• Identify and categorize model failures and safety vulnerabilities.
• Assess bias, misinformation, misuse, and possibly harmful behaviors of the models.
• Detect recurring or systemic issues in model responses.
• Evaluate the severity, reproducibility, and practical significance of identified vulnerabilities.
• Generate high-quality human evaluation data from red teaming efforts.
• Document model failures and classify vulnerabilities accordingly.
• Develop structured attack cases and associated evaluation documentation.
• Ensure consistency across repeated evaluations and datasets.
• Clearly and reproducibly document adversarial scenarios.
• Create reports, datasets, and structured findings for technical teams.
• Communicate identified risks to both technical and non-technical stakeholders.
• Record testing methods, model behaviors, and relevant failure patterns.
• Contribute to broader evaluation efforts across various models and use cases.
• Fluency in both English and Swedish at a native level is mandatory.
• Previous experience in AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing is essential.
• Strong comprehension of conversational AI systems and their failure modes.
• Experience in creating structured adversarial tests and evaluation frameworks.
• Ability to uncover subtle vulnerabilities and recurring behavioral patterns.
• Excellent analytical reasoning and written communication skills.
• Capability to document findings in a clear and reproducible manner.
• Comfortable working on varying projects, scenarios, and evaluation frameworks.
• A background in computer science, cybersecurity, artificial intelligence, machine learning, linguistics, behavioral science, or a related field may be advantageous.
• Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation may also be acceptable.
• Practical experience in red teaming is especially valuable.
• Familiarity with adversarial machine learning concepts.
• Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation topics.
• Experience in cybersecurity including penetration testing, exploit development, or reverse engineering.
• Proven experience in testing conversational AI systems.
• Native-level proficiency in English and Swedish is required for this role.
• This position is available as an independent contractor.
• Work fully remotely.
• Enjoy flexible scheduling.
• Receive competitive hourly pay.
• Get paid weekly via Stripe or Wise.
• Optional participation in higher-sensitivity projects.
• Engage in flexible remote consulting work.
• Project durations may be extended, shortened, or modified based on scope and performance.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.