
Researcher, Evaluations and Benchmarks
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in California.
β’ Deliver a benchmark every two to three weeks that evaluates a previously unmeasured frontier risk.
β’ Partner with top AI laboratories and universities to work on benchmarks and academic papers.
β’ Manage the benchmark taxonomy, harness, quality standards, and release processes.
β’ Personally review evaluations and ensure alignment with taxonomy, rubrics, verifiers, and data distribution.
β’ Oversee the benchmark plan and schedule, ensuring researchers adhere to timelines.
β’ Supervise two or three freelancers/subject matter experts as required.
β’ Meet approximately monthly with the CTO, pod, and research leaders to update the quarterly release plan.
β’ Connect the release roadmap to targeted accounts.
β’ Dedicate roughly 20% of time to monitoring the ecosystem and reviewing research.
β’ Maintain relationships within AI labs and engage with lab personnel weekly.
β’ Travel to conferences a few times each year.
β’ PhD or Master's degree in computer science, machine learning, or a related area, or equivalent experience from industry research.
β’ Over 3 years of experience in building and executing safety or security evaluations for language models in production, within an AI lab, a model provider, or a safety and security research organization.
β’ At least 5 relevant research publications in AI safety and security, including being the lead author on a minimum of 2.
β’ Proficient engineering skills, encompassing evaluation harnesses, distributed inference, vLLM, and the ability to read and modify codebases.
β’ Capability to construct a taxonomy, not just evaluate against one.
β’ Proficiency in guiding a researcher and two freelancers without formal management responsibilities.
β’ Excellent command of English, both written and spoken.
β’ Eagerness to learn about AI harms and adapt to new subjects every three weeks.
β’ Ideally: experience post-training with SFT, DPO, or GRPO.
β’ Ideally: experience with agentic evaluation that includes tool use, orchestration, permissions, and prompt injection.
β’ Ideally: publications in leading conferences.
β’ Willingness to present personal work during client calls.
β’ Strong verbal and written communication skills, with the ability to present to large and/or senior audiences.
β’ Openness to travel to conferences at least 3 times a year.
β’ Budget for engaging freelancers (subject matter experts) as needed.
β’ Travel to conferences at least 3 times annually.
β’ Opportunity to collaborate with prestigious AI labs and universities.
β’ Access to around 150 researchers focused on AI harms.
β’ Conference travel a few times each year / at least 3 times a year.
Orbital Engineering, Inc.
TRIGA Consulting GmbH & Co. KG
Curana Health
SEGUROS INBURSA S.A., GRUPO FINANCIERO INBURSA
Get handpicked remote jobs straight to your inbox weekly.