
AI Safety Red Teamer
Posted Sep 7

Posted Sep 7
This is a fully remote position, open to applicants in United States.
β’ Create adversarial prompts to rigorously test frontier AI models.
β’ Detect jailbreaks, unsafe behaviors, hallucinations, and failures in policy.
β’ Assess model resilience in areas such as misinformation, cyber threats, biosecurity, fraud, political content, and other sensitive fields.
β’ Record vulnerabilities and assist in compiling safety benchmarking and red-teaming reports.
β’ Partner with AI researchers to enhance model alignment, robustness, and safety.
β’ Engage in projects aimed at training and improving AI systems.
β’ A Bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related field.
β’ Over 5 years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a related area.
β’ Excellent analytical reasoning, prompt design, and written communication skills.
β’ Proven experience in designing adversarial prompts or evaluating frontier AI systems.
β’ Preferred experience with AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety.
β’ Familiarity with methodologies related to jailbreak testing, prompt engineering, or adversarial evaluations is preferred.
β’ Expertise in one or more grey-area domains, including cyber, biosecurity, political content, misinformation, or scientific safety is preferred.
β’ Must not require H1-B or STEM OPT sponsorship.
β’ Fully remote position.
β’ Flexible scheduling.
β’ Weekly payments through Stripe or Wise.
β’ Opportunity to collaborate with top AI researchers and safety teams.
β’ Engage in pioneering adversarial testing.
β’ Have a significant impact on AI safety and model development.
β’ Referral bonuses of up to $340 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.