AI Red Teamer – LLM Generalist

atHandshakeRemoteUS flagUnited StatesFreelanceArtificial IntelligenceMid-levelSenior$32 – $95/hour

Posted Sep 15

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Conduct stress tests on large language models by deliberately attempting to break them.

β€’ Create innovative, adversarial prompts that reveal unsafe content, biases, ineffective guardrails, hallucinations, vulnerabilities to prompt injection, and unanticipated behaviors.

β€’ Assess models in various areas including content safety, CBRN, cybersecurity, persuasion and influence operations, child safety, self-harm, over-dependence, and regulatory compliance.

β€’ Evaluate text, image, voice, and agentic model capabilities as project requirements dictate.

β€’ Develop multi-turn scenarios to rigorously test AI safety measures.

β€’ Identify ways to bypass safety filters, restrictions, and defenses using techniques such as jailbreak, evasion, and prompt injection.

β€’ Investigate edge cases to elicit disallowed, harmful, or inaccurate outputs.

β€’ Assess and score model responses using structured harm taxonomies and severity scales.

β€’ Document experiments meticulously, detailing methods, reasoning, and results.

β€’ Review and enhance adversarial prompts created by team members.

β€’ Contribute to the advancement of harm taxonomy, calibration exercises, and inter-rater reliability initiatives.

β€’ Collaborate with engineers, data scientists, and researchers to disseminate findings and bolster defenses.

β€’ Regularly handle potentially distressing content.

β€’ Keep abreast of the latest jailbreaks, attack strategies, and evolving model behaviors.

β€’ Support Handshake AI's collaboration with premier AI research laboratories to enhance model safety and resilience.


⛳️ Requirements

β€’ Extensive hands-on experience with various LLMs, including ChatGPT, Claude, Gemini, and open-source models.

β€’ An innate ability to craft adversarial prompts; knowledge of jailbreak or evasion techniques is a significant advantage.

β€’ Creative and adversarial problem-solving abilities.

β€’ Clear and thoughtful written communication skills.

β€’ Strong ethical judgment and the capacity to distinguish adversarial thinking from personal morals.

β€’ Self-motivated, collaborative, and at ease in feedback-rich environments.

β€’ A sense of curiosity, persistence, and comfort with frequent setbacks in experimentation.

β€’ Candidates must be capable of professionally and sustainably engaging with harmful material.

β€’ Familiarity with Python or other scripting languages.

β€’ Experience with LLM APIs or evaluation tools.

β€’ Competence in structured data annotation and rubric-based scoring.

β€’ Previous work in trust and safety, content moderation, QA, or security research.

β€’ Expertise in a high-risk field such as cybersecurity, chemistry, biology, medicine, law, or finance.

β€’ Ability to work remotely from the United States, Monday through Friday, for 40 hours per week.

β€’ Must be authorized to work lawfully in the United States for Handshake.

β€’ Must not require employment visa sponsorship, as specified in the application guidelines.


🏝️ Benefits

β€’ Support resources are available for managing exposure to disturbing content.

β€’ Cash compensation ranges from $32 to $95 per hour.

β€’ Opportunity for remote work from the United States.

β€’ Full-time commitment of 40 hours per week.

People also viewed

Mercor14 hours ago

AI Safety Red Teamer

US flagUnited States OnlyFreelanceArtificial Intelligence$70 – $84/hour
ApplyView job
Mercor15 hours ago

AI Safety Expert – English, Telugu

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor15 hours ago

AI Safety Experts, English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor1 day ago

AI Safety Expert, English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor1 day ago

AI Safety Experts – English, Punjabi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor1 day ago

AI Safety Expert – English, Gujarati

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers