
Senior AI Backend Engineer – Agent Evaluation, Quality
Posted Jul 19

Posted Jul 19
This is a fully remote position, open to applicants in Saudi Arabia.
• Take charge of the evaluation framework. Design and develop LLM-as-judge systems, calibrate them using human labels, and create measurable quality metrics for agents and specific failure modes.
• Implement the release gate effectively. Create evaluation harnesses for each pull request and integrate regression detection into continuous integration, ensuring quality is maintained automatically rather than relying on manual checks.
• Develop user simulators to produce test coverage and adversarial scenarios before they reach actual users.
• Convert production signals into enhancements - funnel real failures back into evaluation sets to enable the system to improve over time.
• Collaborate with product teams to define "what good looks like" through clear, measurable criteria.
• Expand into agent development - assist in constructing and refining the agents themselves, beginning with the aspects you are most familiar with from your evaluation experiences.
• Solid foundation in software engineering principles. Proficient in production-level Python or Typescript (or similar), clean API and system design, testing, and CI/CD. You write code that others can build upon - evaluation infrastructure is a key aspect of real engineering.
• Practical experience with LLMs and agents. You have developed applications using LLMs - including agents, RAG, tool/function invocation, and orchestration frameworks (such as LangGraph, LangChain, or similar) - and have insights into their behavior and potential failures.
• A metrics-oriented mindset. You analyze metrics, calibration, and experiments; you aim to quantify the effectiveness of a solution rather than simply deploying it.
• Experience in production environments. You have operated LLM systems in live settings and have managed aspects like reliability, latency, cost, and observability.
• 5+ years of software engineering experience, with recent involvement in LLM/agent projects.
• Nice to have
• Direct experience in evaluating LLM and agent systems - including offline/online evaluation, LLM-as-judge, and systematic regression testing.
• Familiarity with observability tools (such as Arize, LangSmith, or similar).
• Knowledge of Arabic language and NLP.
• Experience with e-commerce or merchant-facing products.
Public Partnerships | PPL
Winkler & Company
Centene Corporation
Robots & Pencils
Get handpicked remote jobs straight to your inbox weekly.