
AI Red Teamer β LLM Generalist
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
β’ Conduct stress tests on large language models by deliberately attempting to break them.
β’ Create innovative, adversarial prompts that reveal unsafe content, biases, ineffective guardrails, hallucinations, vulnerabilities to prompt injection, and unanticipated behaviors.
β’ Assess models in various areas including content safety, CBRN, cybersecurity, persuasion and influence operations, child safety, self-harm, over-dependence, and regulatory compliance.
β’ Evaluate text, image, voice, and agentic model capabilities as project requirements dictate.
β’ Develop multi-turn scenarios to rigorously test AI safety measures.
β’ Identify ways to bypass safety filters, restrictions, and defenses using techniques such as jailbreak, evasion, and prompt injection.
β’ Investigate edge cases to elicit disallowed, harmful, or inaccurate outputs.
β’ Assess and score model responses using structured harm taxonomies and severity scales.
β’ Document experiments meticulously, detailing methods, reasoning, and results.
β’ Review and enhance adversarial prompts created by team members.
β’ Contribute to the advancement of harm taxonomy, calibration exercises, and inter-rater reliability initiatives.
β’ Collaborate with engineers, data scientists, and researchers to disseminate findings and bolster defenses.
β’ Regularly handle potentially distressing content.
β’ Keep abreast of the latest jailbreaks, attack strategies, and evolving model behaviors.
β’ Support Handshake AI's collaboration with premier AI research laboratories to enhance model safety and resilience.
β’ Extensive hands-on experience with various LLMs, including ChatGPT, Claude, Gemini, and open-source models.
β’ An innate ability to craft adversarial prompts; knowledge of jailbreak or evasion techniques is a significant advantage.
β’ Creative and adversarial problem-solving abilities.
β’ Clear and thoughtful written communication skills.
β’ Strong ethical judgment and the capacity to distinguish adversarial thinking from personal morals.
β’ Self-motivated, collaborative, and at ease in feedback-rich environments.
β’ A sense of curiosity, persistence, and comfort with frequent setbacks in experimentation.
β’ Candidates must be capable of professionally and sustainably engaging with harmful material.
β’ Familiarity with Python or other scripting languages.
β’ Experience with LLM APIs or evaluation tools.
β’ Competence in structured data annotation and rubric-based scoring.
β’ Previous work in trust and safety, content moderation, QA, or security research.
β’ Expertise in a high-risk field such as cybersecurity, chemistry, biology, medicine, law, or finance.
β’ Ability to work remotely from the United States, Monday through Friday, for 40 hours per week.
β’ Must be authorized to work lawfully in the United States for Handshake.
β’ Must not require employment visa sponsorship, as specified in the application guidelines.
β’ Support resources are available for managing exposure to disturbing content.
β’ Cash compensation ranges from $32 to $95 per hour.
β’ Opportunity for remote work from the United States.
β’ Full-time commitment of 40 hours per week.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.