
AI Safety Red Teamer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
β’ Create adversarial prompts to rigorously test advanced AI models.
β’ Detect jailbreaks, unsafe behaviors, hallucinations, and failures in policy implementation.
β’ Assess model resilience in sensitive areas such as misinformation, cybersecurity, biosecurity, fraud, and political content.
β’ Record vulnerabilities and assist in the preparation of safety benchmarking and red-teaming documentation.
β’ Work alongside AI researchers to enhance model alignment, robustness, and safety measures.
β’ Conduct adversarial testing to train and improve next-generation AI systems.
β’ A Bachelor's degree or higher in fields such as Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related area.
β’ Over 5 years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or similar fields.
β’ Excellent analytical reasoning abilities.
β’ Proficient skills in prompt design.
β’ Strong written communication capabilities.
β’ Experience in designing adversarial prompts or assessing advanced AI systems.
β’ Preferred experience in AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety.
β’ Familiarity with methodologies related to jailbreak testing, prompt engineering, or adversarial evaluation is preferred.
β’ Expertise in one or more grey-area domains, such as cybersecurity, biosecurity, political content, misinformation, or scientific safety is preferred.
β’ Ability to work independently as a contractor is essential.
β’ H1-B and STEM OPT candidates will not be supported.
β’ Fully remote work environment.
β’ Flexible scheduling options.
β’ Project timelines can be adjusted or concluded early based on needs and performance.
β’ Weekly payments through Stripe or Wise based on services provided.
β’ Competitive compensation.
β’ Opportunity to collaborate with leading AI researchers and safety teams.
β’ Chance to engage in innovative adversarial testing.
β’ Influence the safety and development of frontier AI models.
β’ Referral bonuses of up to $340 for each successful referral.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.