Remotery

Senior AI Backend Engineer – Agent Evaluation, Quality

Posted Jul 19

This is a fully remote position, open to applicants in Saudi Arabia.

📋 Description

• Take charge of the evaluation framework. Design and develop LLM-as-judge systems, calibrate them using human labels, and create measurable quality metrics for agents and specific failure modes.

• Implement the release gate effectively. Create evaluation harnesses for each pull request and integrate regression detection into continuous integration, ensuring quality is maintained automatically rather than relying on manual checks.

• Develop user simulators to produce test coverage and adversarial scenarios before they reach actual users.

• Convert production signals into enhancements - funnel real failures back into evaluation sets to enable the system to improve over time.

• Collaborate with product teams to define "what good looks like" through clear, measurable criteria.

• Expand into agent development - assist in constructing and refining the agents themselves, beginning with the aspects you are most familiar with from your evaluation experiences.


⛳️ Requirements

• Solid foundation in software engineering principles. Proficient in production-level Python or Typescript (or similar), clean API and system design, testing, and CI/CD. You write code that others can build upon - evaluation infrastructure is a key aspect of real engineering.

• Practical experience with LLMs and agents. You have developed applications using LLMs - including agents, RAG, tool/function invocation, and orchestration frameworks (such as LangGraph, LangChain, or similar) - and have insights into their behavior and potential failures.

• A metrics-oriented mindset. You analyze metrics, calibration, and experiments; you aim to quantify the effectiveness of a solution rather than simply deploying it.

• Experience in production environments. You have operated LLM systems in live settings and have managed aspects like reliability, latency, cost, and observability.

• 5+ years of software engineering experience, with recent involvement in LLM/agent projects.

• Nice to have

• Direct experience in evaluating LLM and agent systems - including offline/online evaluation, LLM-as-judge, and systematic regression testing.

• Familiarity with observability tools (such as Arize, LangSmith, or similar).

• Knowledge of Arabic language and NLP.

• Experience with e-commerce or merchant-facing products.

People also viewed

Public Partnerships | PPL2 days ago

QA Engineer – AI

US flagUnited States OnlyFull-timeQA Engineer (Quality Assurance)$75k – $80k/year
ApplyView job
Winkler & Company2 days ago

App Tester, Online Test User

DE flagGermany OnlyPart-timeQA Engineer (Quality Assurance)
ApplyView job
Centene Corporation2 days ago

Senior Manager, Quality Assurance

US flagMissouri OnlyFull-timeQA Engineer (Quality Assurance)$102.9k – $190.5k/year
ApplyView job
Robots & Pencils2 days ago

Senior Testing Engineer

CO flagColombia OnlyFull-timeQA Engineer (Quality Assurance)
ApplyView job
Regeneron2 days ago

Manager, Quality Assurance Auditor – GCP

US flagUnited States OnlyFull-timeQA Engineer (Quality Assurance)$114.8k – $187.4k/year
ApplyView job
Nectir2 days ago

Quality Assurance Engineer

US flagUnited States, +1 more stateFull-timeQA Engineer (Quality Assurance)$54k – $152k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers