
AI Safety Specialist, English, Norwegian
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in New York.
• Perform adversarial testing on conversational AI models and agents.
• Design jailbreaks, prompt-injection scenarios, misuse cases, and multi-turn manipulation tactics.
• Examine models for vulnerabilities in various conversational and adversarial contexts.
• Detect and categorize model failures and safety vulnerabilities.
• Analyze bias, misinformation, misuse, and potentially dangerous model behaviors.
• Evaluate the severity, reproducibility, and practical implications of vulnerabilities.
• Generate and annotate human evaluation data from red teaming exercises.
• Develop structured attack cases and assessment materials.
• Document adversarial scenarios, testing methodologies, model behaviors, and failure patterns.
• Create reports, datasets, and organized findings for technical teams.
• Communicate identified risks to both technical and non-technical stakeholders.
• Support broader evaluation efforts across different models and use cases.
• Native-level proficiency in both English and Norwegian.
• Prior experience in AI red teaming, adversarial AI assessment, cybersecurity, or socio-technical system testing.
• Strong grasp of conversational AI systems and their failure modes.
• Experience in developing structured adversarial tests.
• Capability to identify subtle vulnerabilities and recurring behavioral patterns.
• Excellent analytical reasoning and written communication abilities.
• Proficient in documenting findings clearly and reproducibly.
• Comfortable working across varying projects, scenarios, and evaluation frameworks.
• Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation may be accepted.
• Practical experience in red teaming is especially valuable.
• Familiarity with adversarial machine learning concepts.
• Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation concepts.
• Cybersecurity experience in penetration testing, exploit development, or reverse engineering.
• Experience analyzing abuse, harassment, misinformation, or other socio-technical risks.
• Background in testing conversational AI systems.
• Previous experience in generating structured human data for AI evaluation.
• Fully remote work environment.
• Flexible scheduling options.
• Competitive hourly pay.
• Weekly payments via Stripe or Wise.
• Optional participation in higher-sensitivity projects.
• Flexible remote consulting opportunities.
• Project timelines may be extended, shortened, or adjusted based on scope and performance.
Progressive Leasing
apna
apna
Texas Research International
Get handpicked remote jobs straight to your inbox weekly.