
AI Safety Red Teamer
Posted 3 hours ago

Posted 3 hours ago
This is a fully remote position, open to applicants in United States.
β’ Develop adversarial prompts to rigorously test frontier AI models.
β’ Detect jailbreaks, unsafe behaviors, hallucinations, and failures in policy.
β’ Assess model resilience in areas such as misinformation, cybersecurity, biosecurity, fraud, political content, and other sensitive topics.
β’ Record vulnerabilities and aid in the creation of safety benchmarking and red-teaming documentation.
β’ Work in collaboration with AI researchers to enhance model alignment, robustness, and safety.
β’ Engage in project activities for Mercor's partner AI labs and enterprises to train and improve frontier AI systems.
β’ A bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related field.
β’ Over 5 years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a comparable area.
β’ Excellent analytical reasoning, prompt design, and written communication abilities.
β’ Experience in creating adversarial prompts or assessing frontier AI systems.
β’ Preferred experience in AI Red Teaming, Reinforcement Learning from Human Feedback (RLHF), Supervised Fine-Tuning (SFT), AI Alignment, or Trust & Safety.
β’ Familiarity with jailbreak testing, prompt engineering, or methodologies for adversarial evaluation is preferred.
β’ Expertise in ambiguous domains such as cybersecurity, biosecurity, political content, misinformation, or scientific safety is preferred.
β’ Must be eligible to work in one of the specified location-requirement countries.
β’ H1-B and STEM OPT candidates are not supported.
β’ Completely remote position.
β’ Flexible work schedule; tasks can be completed at your convenience.
β’ Weekly payments via Stripe or Wise.
β’ Chance to engage in groundbreaking adversarial testing with top AI researchers and safety teams.
β’ Opportunity to shape how AI systems address intricate, real-world safety issues.
β’ Competitive salary.
β’ Referral bonuses of up to $340 for each successful referral, subject to referral limits.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.