
AI Safety Red Teamer
Posted 4 hours ago

Posted 4 hours ago
This is a fully remote position, open to applicants in United States.
β’ Create adversarial prompts to evaluate and challenge advanced AI models
β’ Detect jailbreaks, unsafe behaviors, hallucinations, and policy breaches
β’ Assess model resilience in areas such as misinformation, cybersecurity, biosecurity, fraud, political content, and other sensitive sectors
β’ Record vulnerabilities and assist in safety benchmarking and red-teaming documentation
β’ Partner with AI researchers to enhance model alignment, robustness, and safety measures
β’ Engage in projects aimed at training and refining AI systems
β’ A Bachelor's degree or higher in fields such as Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related area
β’ Over 5 years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a comparable field
β’ Excellent analytical reasoning, prompt design, and written communication abilities
β’ Experience in designing adversarial prompts or assessing advanced AI systems
β’ Preferred: background in AI Red Teaming, RLHF, SFT, AI Alignment, or Trust & Safety
β’ Preferred: knowledge of jailbreak testing, prompt engineering, or adversarial evaluation techniques
β’ Preferred: expertise in one or more unconventional areas, such as cybersecurity, biosecurity, political content, misinformation, or scientific safety
β’ Must be capable of working as an independent contractor
β’ H1-B and STEM OPT candidates are not eligible
β’ Fully remote position
β’ Flexible working hours; tasks can be completed according to your own schedule
β’ Weekly payments through Stripe or Wise based on services provided
β’ Project durations can be extended, shortened, or concluded early based on requirements and performance
β’ Chance to collaborate on innovative adversarial testing projects with top AI researchers and safety teams
β’ Opportunity to shape how AI systems address intricate, real-world safety challenges
β’ Competitive compensation
β’ Referral bonus of up to $340 for each successful referral
XenoPatch GmbH
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.