Remotery

LLM Red Team & Benchmark Evaluation Specialist

at24-MAGRemoteUS flagNew YorkFull-timeUncategorizedJunior$55 – $85/hour

Posted Jul 27

This is a fully remote position, open to applicants in New York.

📋 Description

• Evaluation of adversarial models

• Test cutting-edge AI models in areas such as coding, machine learning, analysis, and multi-step agentic tasks

• Detect subtle errors, vulnerabilities, edge cases, and outputs that may appear plausible but are misleading

• Explore scenarios where models seem competent but arrive at incorrect or unsupported conclusions

• Create reproducible experiments to identify and confirm model failure modes

• Benchmarking & Challenge Creation

• Transform identified weaknesses in models into rigorous benchmarking tasks

• Design challenges that are technically rigorous while remaining fair and objectively assessable

• Clearly define task requirements, expected results, and evaluation metrics

• Ensure that tasks necessitate genuine reasoning instead of permitting shortcuts or superficial pattern matching

• Failure Analysis & Reporting

• Document findings with solid evidence, methodologies, and reproducible processes

• Clarify the reasons behind model failures and the capabilities or assumptions that led to the errors

• Generate comprehensive technical write-ups for researchers and task creators

• Monitor recurring failure patterns across different models, prompts, and evaluation environments

• Task Enhancement & Collaborative Research

• Collaborate with task authors to address loopholes, grading inconsistencies, and unintended shortcuts

• Review benchmark tasks for clarity, potential for exploitation, and evaluation reliability

• Share insights with researchers and specialists to enhance benchmark coverage

• Engage in iterative calibration, peer review, and task refinement processes


⛳️ Requirements

• Minimum of 1 year of experience in research, research engineering, security, AI evaluation, or a related technical field

• Proven experience in identifying vulnerabilities, edge cases, or failure modes in LLMs or ML systems

• Background in red teaming, adversarial testing, security research, benchmark development, or thorough model evaluation

• Proficient in Python and Git

• Capability to independently develop scripts, probes, and analyses

• Strong understanding of LLM capabilities, limitations, and evaluation methodologies

• Exceptional written communication and technical documentation abilities

• Creativity, precision, and persistence in addressing ambiguous research challenges

• Reliable availability for approximately 35 hours per week

• A master's degree or PhD in a STEM field is highly relevant


🏝️ Benefits

• Opportunity to work on cutting-edge AI technology

• Collaborative and innovative work environment

• Professional development and growth opportunities

People also viewed

Behavioral Health Works, Inc.7 hours ago

BCBA

US flagTexas OnlyFull-timeUncategorized$90k – $110k/year
ApplyView job
Sodexo8 hours ago

Strategic Solutions & Proposal Manager – Food & Facilities Services

CA flagCanada OnlyFull-timeUncategorizedC$110k – C$135k/year
ApplyView job
Sodexo8 hours ago

Strategic Solutions & Proposal Manager – Food & Facilities Services

CA flagCanada OnlyFull-timeUncategorizedC$125k – C$150k/year
ApplyView job
EVERSANA8 hours ago

Specialty Pharmaceutical Representative – Women's Health

US flagArizona OnlyFull-timeUncategorized$90k – $130k/year
ApplyView job
Pearce Services8 hours ago

Generator Technician

US flagTennessee OnlyFull-timeUncategorized$30 – $40/hour
ApplyView job
Winning Assistants LLC8 hours ago

Behavioral Health Patient Coordinator – Charting

PH flagPhilippines OnlyFull-timeUncategorized$5 – $6/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers