
AI Evaluation Engineer
Posted Jul 18

Posted Jul 18
This is a fully remote position, open to applicants in Colorado.
• Create, develop, and sustain automated AI evaluation pipelines for production LLM applications.
• Formulate prompt engineering strategies, refine prompts, and assess LLMs using quantitative evaluation techniques.
• Construct offline evaluation datasets and regression testing frameworks to track AI performance over time.
• Examine production AI behavior utilizing Python, SQL, and statistical methods to pinpoint areas for enhancement.
• Design experiments, A/B tests, and benchmarking methods for assessing prompt and model modifications.
• Create dashboards and reports that convey AI quality, reliability, and performance metrics.
• Collaborate with engineering and product teams to safely implement and oversee enhancements to production AI systems.
• Explore model failures through thorough error analysis and propose enhancements to prompts, evaluation datasets, and workflows.
• Assist in establishing best practices for Responsible AI, evaluation methodologies, and ongoing model enhancement.
• Must possess US citizenship and successfully pass an FBI fingerprint and background check across multiple states.
• 3+ years of experience in software engineering, machine learning, data science, or a comparable technical discipline.
• Proficient in designing evaluation metrics and analyzing AI model performance.
• Knowledge of statistical methods, including hypothesis testing and experimental design.
• Strong experience in Python development.
• Solid SQL skills with experience in analyzing extensive datasets.
• Background in building or supporting production LLM or Generative AI applications.
• Familiarity with prompt engineering and systematic evaluation of prompts.
• Offers bonus opportunities.
• Health, dental, and vision benefits.
• Flexible time off policy.
Quantiphi
BJAK
BJAK
Get handpicked remote jobs straight to your inbox weekly.