
LLM Red Team & Benchmark Evaluation Specialist
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in New York.
• Evaluation of adversarial models
• Test cutting-edge AI models in areas such as coding, machine learning, analysis, and multi-step agentic tasks
• Detect subtle errors, vulnerabilities, edge cases, and outputs that may appear plausible but are misleading
• Explore scenarios where models seem competent but arrive at incorrect or unsupported conclusions
• Create reproducible experiments to identify and confirm model failure modes
• Benchmarking & Challenge Creation
• Transform identified weaknesses in models into rigorous benchmarking tasks
• Design challenges that are technically rigorous while remaining fair and objectively assessable
• Clearly define task requirements, expected results, and evaluation metrics
• Ensure that tasks necessitate genuine reasoning instead of permitting shortcuts or superficial pattern matching
• Failure Analysis & Reporting
• Document findings with solid evidence, methodologies, and reproducible processes
• Clarify the reasons behind model failures and the capabilities or assumptions that led to the errors
• Generate comprehensive technical write-ups for researchers and task creators
• Monitor recurring failure patterns across different models, prompts, and evaluation environments
• Task Enhancement & Collaborative Research
• Collaborate with task authors to address loopholes, grading inconsistencies, and unintended shortcuts
• Review benchmark tasks for clarity, potential for exploitation, and evaluation reliability
• Share insights with researchers and specialists to enhance benchmark coverage
• Engage in iterative calibration, peer review, and task refinement processes
• Minimum of 1 year of experience in research, research engineering, security, AI evaluation, or a related technical field
• Proven experience in identifying vulnerabilities, edge cases, or failure modes in LLMs or ML systems
• Background in red teaming, adversarial testing, security research, benchmark development, or thorough model evaluation
• Proficient in Python and Git
• Capability to independently develop scripts, probes, and analyses
• Strong understanding of LLM capabilities, limitations, and evaluation methodologies
• Exceptional written communication and technical documentation abilities
• Creativity, precision, and persistence in addressing ambiguous research challenges
• Reliable availability for approximately 35 hours per week
• A master's degree or PhD in a STEM field is highly relevant
• Opportunity to work on cutting-edge AI technology
• Collaborative and innovative work environment
• Professional development and growth opportunities
Behavioral Health Works, Inc.
Sodexo
Sodexo
EVERSANA
Get handpicked remote jobs straight to your inbox weekly.