
Senior Data Scientist – AI Evaluation
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in North America.
• Create AI evaluation frameworks by establishing ground truth, metrics, and scoring techniques for models and agents.
• Develop repeatable evaluation processes to monitor quality over time and identify regressions prior to deployment.
• Collaborate with engineering and analytics engineering teams to implement evaluation harnesses.
• Convert evaluation outcomes into actionable insights for system enhancements.
• Set up evaluation protocols, documentation, and review processes.
• Guide and harmonize teams on evaluation best practices and quantifiable AI quality.
• Collaborate with Product, Engineering, Analytics Engineering, and business stakeholders to establish quality benchmarks and encourage iteration.
• Maintain the quality standards independently from the teams responsible for building and optimizing the systems.
• Proven experience in quantitative measurement rigor (e.g., LLM/model evaluation, metric validation, or experimentation).
• Strong foundation in statistics and machine learning, including sample sizing, confidence intervals, significance testing, managing non-determinism, and validating automated grading against human ground truth.
• Proficiency in Python and SQL, with experience in evaluating models within production settings.
• Good judgment in defining quality metrics and establishing ground truth for ambiguous outputs.
• Exceptional communication and cross-functional collaboration skills.
• Strong problem-solving skills in dynamic, greenfield environments.
• 6–10 years of experience in quantitative data science or machine learning, with specialized experience in measurement or evaluation.
• A quantitative degree is advantageous; equivalent experience in the industry is also acceptable.
• Nice to have: direct experience with LLM/agent evaluation in production, evaluation harnesses, LLM-as-judge calibration, and continuous integration regression gates.
• Nice to have: experience in evaluating text-to-SQL, analytics agents, or systems where correctness can be verified against data.
• Nice to have: background in fintech, brokerage, or sectors with business or risk implications.
• Nice to have: familiarity with AI tools in research and engineering workflows.
• Competitive Salary & Stock Options
• Health Benefits
• New Hire Home-Office Setup: One-time USD $500
• Monthly Stipend: USD $150 per month via a Brex Card
Tendios
GuidePoint Security
Seneca Holdings
Get handpicked remote jobs straight to your inbox weekly.