AI Safety Red Teamer

atMercorRemoteUS flagUnited StatesFreelanceArtificial IntelligenceMid-levelSenior$70 – $84/hour

Posted 3 hours ago

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Develop adversarial prompts to rigorously test frontier AI models.

β€’ Detect jailbreaks, unsafe behaviors, hallucinations, and failures in policy.

β€’ Assess model resilience in areas such as misinformation, cybersecurity, biosecurity, fraud, political content, and other sensitive topics.

β€’ Record vulnerabilities and aid in the creation of safety benchmarking and red-teaming documentation.

β€’ Work in collaboration with AI researchers to enhance model alignment, robustness, and safety.

β€’ Engage in project activities for Mercor's partner AI labs and enterprises to train and improve frontier AI systems.


⛳️ Requirements

β€’ A bachelor's degree or higher in Computer Science, Cybersecurity, Journalism, Communications, Psychology, Biology, Chemistry, Public Policy, or a related field.

β€’ Over 5 years of professional experience in AI Safety, AI Red Teaming, Trust & Safety, cybersecurity, investigative journalism, life sciences, or a comparable area.

β€’ Excellent analytical reasoning, prompt design, and written communication abilities.

β€’ Experience in creating adversarial prompts or assessing frontier AI systems.

β€’ Preferred experience in AI Red Teaming, Reinforcement Learning from Human Feedback (RLHF), Supervised Fine-Tuning (SFT), AI Alignment, or Trust & Safety.

β€’ Familiarity with jailbreak testing, prompt engineering, or methodologies for adversarial evaluation is preferred.

β€’ Expertise in ambiguous domains such as cybersecurity, biosecurity, political content, misinformation, or scientific safety is preferred.

β€’ Must be eligible to work in one of the specified location-requirement countries.

β€’ H1-B and STEM OPT candidates are not supported.


🏝️ Benefits

β€’ Completely remote position.

β€’ Flexible work schedule; tasks can be completed at your convenience.

β€’ Weekly payments via Stripe or Wise.

β€’ Chance to engage in groundbreaking adversarial testing with top AI researchers and safety teams.

β€’ Opportunity to shape how AI systems address intricate, real-world safety issues.

β€’ Competitive salary.

β€’ Referral bonuses of up to $340 for each successful referral, subject to referral limits.

People also viewed

Mercor4 hours ago

AI Safety Expert – English, Telugu

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor4 hours ago

AI Safety Experts, English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor13 hours ago

AI Safety Expert, English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor13 hours ago

AI Safety Experts – English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor13 hours ago

AI Safety Expert – English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Aston Carter13 hours ago

Life Sciences AI Trainer

US flagMassachusetts OnlyFreelanceArtificial Intelligence$75/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers