
Software Engineering Director, Agentic Evaluations
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in United States.
• Assess the feasibility and effort required for evaluating software-vendor agents.
• Manage the complete stack of the core evaluation system, which includes administrative and API interfaces as well as the workflow for evaluations.
• Create reusable components for the evaluation system across different software domains.
• Transform integration onboarding and maintenance tasks into repeatable AI capabilities or agents.
• Keep track of new frameworks and methodologies for agent evaluation.
• Set up industry-standard practices for disseminating evaluation outcomes.
• Partner with data science colleagues to establish proprietary benchmarks.
• Provide leadership, mentorship, and technical guidance to a core engineering team.
• Advocate for the incorporation of evaluations in agent-focused products throughout the organization.
• Produce valuable and trustworthy evaluations of active software-vendor agents.
• Utilize agent-first engineering strategies to deliver work for production.
• Over 10 years of professional programming experience in backend or full-stack environments.
• More than 2 years of experience in directly managing engineering teams.
• High-level backend development expertise in Python, Java/Kotlin, TypeScript/JavaScript, or Go.
• Strong skills in backend frameworks such as FastAPI or Node.js.
• Hands-on experience creating evaluations for customer-facing agents using trajectory trace data and evaluation rubrics.
• Experience in measuring metrics like task completion rates, accuracy, correctness, or policy compliance.
• Direct experience with frontier models from OpenAI, Anthropic, or Google in LLM-as-a-judge applications.
• Regular use of coding agent harnesses such as Claude Code, Codex, Opencode, or Pi.
• Bachelor’s degree in a related field.
• Experience with agent tool usage through direct integration or MCP servers is a plus.
• Familiarity with benchmark frameworks like STATE-Bench, tau2-bench, or similar is advantageous.
• Experience with tools like Playwright, browser usage, Chrome DevTools MCP, or similar is a plus.
• Experience with durable execution/workflow frameworks or agent sandboxing is beneficial.
• Authorized to work in the United States without restrictions.
• No current or future requirement for U.S. work sponsorship.
• Equity.
• Bonus opportunities.
• Flexible working arrangements.
• Generous parental leave.
• Unlimited paid time off (PTO).
• An inclusive and diverse work environment.
• Employee resource groups (ERGs).
• G2 Gives philanthropic program.
• Option to opt-out of AI-assisted application reviews.
RoomPriceGenie
Syniti
Stone & Company
CVS Health
Get handpicked remote jobs straight to your inbox weekly.