
AI Safety Specialist, English – Danish
Posted Aug 7

Posted Aug 7
This is a fully remote position, open to applicants in New York.
• Perform adversarial assessments of conversational AI models and agents.
• Create jailbreaks, prompt-injection scenarios, misuse instances, and multi-turn manipulation techniques.
• Examine models for vulnerabilities in various conversational and adversarial contexts.
• Implement systematic testing frameworks.
• Detect and categorize model failures and safety vulnerabilities.
• Analyze bias, misinformation, misuse, and potentially harmful behaviors of models.
• Evaluate the severity, reproducibility, and practical significance of vulnerabilities.
• Generate high-quality human evaluation data from red teaming efforts.
• Annotate model failures and classify identified vulnerabilities.
• Develop structured attack cases and evaluation resources.
• Clearly and reproducibly document adversarial scenarios.
• Produce reports, datasets, and organized findings for technical teams.
• Communicate identified risks to both technical and non-technical stakeholders.
• Record testing methodologies, model behaviors, and failure patterns.
• Contribute to comprehensive evaluation coverage across models and use cases.
• Native-level proficiency in both English and Danish.
• Previous experience in AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing.
• Strong grasp of conversational AI systems and their failure modes.
• Experience in developing structured adversarial tests and evaluation frameworks.
• Ability to identify subtle vulnerabilities and recurring behavioral patterns.
• Excellent analytical reasoning and written communication skills.
• Capability to document findings in a clear and reproducible manner.
• Comfort in working across varying projects, scenarios, and evaluation frameworks.
• A background in computer science, cybersecurity, artificial intelligence, machine learning, linguistics, behavioral science, or a related field may be advantageous.
• Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation may also be accepted.
• Experience in creating jailbreak or prompt-injection datasets.
• Familiarity with adversarial machine learning.
• Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation concepts.
• Cybersecurity experience involving penetration testing, exploit development, or reverse engineering.
• Background in analyzing abuse, harassment, misinformation, or other socio-technical risks.
• Experience in testing conversational AI systems.
• Previous experience in producing structured human data for AI evaluation.
• Native-level English and Danish proficiency is essential.
• Status as an independent contractor is required.
• Flexible scheduling options.
• Competitive hourly pay rate ($45–$60 per hour).
• Weekly payments via Stripe or Wise.
• Optional involvement in higher-sensitivity projects.
• Flexible remote consulting opportunities.
• Project timelines may be extended, shortened, or modified based on scope and performance.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.