
Researcher, Evaluations and Benchmarks
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in New York.
• Deliver a benchmark every two to three weeks that assesses a previously unmeasured frontier risk.
• Collaborate with top AI laboratories and academic institutions on benchmarks and research publications.
• Work alongside in-house researchers responsible for specific harm areas and guide freelancers.
• Manage the taxonomy, evaluation harness, quality standards, and release process.
• Ensure that frontier labs can replicate the benchmark and validate the reported figures.
• Review evaluations, confirm taxonomy alignment, and encourage researchers to maintain high quality.
• Oversee the benchmark strategy and timeline, keeping researchers on track.
• Supervise two or three freelance subject matter experts as necessary.
• Meet approximately once a month with the CTO, pod, and research leads to establish the quarterly release roadmap.
• Align the release plan with target accounts.
• Monitor the AI safety and security landscape, review research, maintain relationships with labs, and collect feedback.
• Attend conferences and engage with AI lab contacts on a weekly basis.
• PhD or Master's degree in computer science, machine learning, or a related discipline, or equivalent experience from industry research.
• More than 3 years of experience in developing and executing safety or security evaluations for language models in production, within an AI lab, model provider, or a safety and security research organization.
• At least 5 pertinent research publications in AI safety and security, including being the lead author on a minimum of 2.
• Strong engineering capabilities, including experience with evaluation harnesses, distributed inference, vLLM, and the ability to read and troubleshoot a codebase.
• Skill in developing a taxonomy, not just scoring against one.
• Capability to direct a researcher and two freelancers without formal management authority.
• Proficient in English, both written and spoken.
• Genuine curiosity about harms and the ability to learn a new subject every three weeks.
• Experience with post-training methods such as SFT, DPO, or GRPO (preferably).
• Experience in designing rewards for subjective and safety-critical objectives (preferably).
• Background in agentic evaluation involving tool use, orchestration, permissions, and prompt injection (preferably).
• Publications presented at leading conferences (preferably).
• Willingness to present work during client calls; excellent verbal and written communication skills and the ability to present to large and/or senior audiences (preferably).
• Budget allocated for freelancers as needed on an ad hoc basis.
• Opportunity to travel to conferences a couple of times each year.
• Ideally, travel to conferences at least three times a year.
ALB Conciergerie
Meiks Affiliate Tipps
StanMindsetMomentum
LEARN Behavioral
Get handpicked remote jobs straight to your inbox weekly.