
Task Development Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants anywhere in the world.
• Create innovative and challenging tasks for AI models that remain engaging as the model's time horizons expand.
• Conduct quality assurance on current tasks, ensuring they are solvable and that models receive only the essential information.
• Establish baseline tasks within your areas of expertise when beneficial.
• Evaluate task completions from AI systems or human baseliners.
• Enhance the infrastructure and workflows for task development.
• Participate in assessments of cutting-edge AI systems in alignment with METR’s Time Horizons methodology.
• Multiple years of experience in complex software engineering projects and codebases.
• Proven experience in developing challenging AI evaluations, preferably with agent-based approaches.
• Familiarity with evaluation frameworks such as RE-Bench, HCAST, SWE-bench Verified, Cybench, or GPQA.
• Experience using the Inspect framework is preferred.
• Strong attention to detail, with the ability to identify misspecifications and ambiguities.
• Knowledge of METR infrastructure, including Hawk, is an advantage.
• Understanding of the methodology behind METR’s Time Horizons initiative is a plus.
• Availability for 20–40 hours per week.
• At least 1 hour of overlap with the Pacific Coast Time workday.
• Schedule flexibility determined by you.
• Opportunity for remote work from anywhere in the world.
• Flexible working hours ranging from 20 to 40 hours per week.
• Contributors of over 80 hours will receive acknowledgment in the final research output (if desired).
LiteLLM AI Gateway
Snowflake
RTX
C-MORE
Get handpicked remote jobs straight to your inbox weekly.