
Senior Data Scientist – AI Evaluation, Improvement
Posted 5 hours ago

Posted 5 hours ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the diagnostic loop for clinical services based on LLM technology.
• Develop root-cause hypotheses for identified failure modes related to prompts, context, retrieval, guidelines, model behavior, and upstream data.
• Create experiments aimed at isolating various causes.
• Evaluate prompt, configuration, context, and model variations.
• Confirm enhancements through structured evaluations while ensuring there are no regressions.
• Manage the evaluation roadmap and maintain quality standards for AI services.
• Construct golden sets and regression suites.
• Define and analyze metrics such as accuracy, adherence to guidelines, and grounding/faithfulness.
• Thoughtfully expand the team's evaluation platform.
• Collaborate with clinical reviewers to convert findings into labeled evidence and evaluation criteria.
• Work alongside product managers to prioritize critical failure modes.
• Mentor colleagues and contribute to the growth of the evaluations function.
• Assist in the annual update cycle of clinical guidelines through regression evaluation.
• Document failure taxonomies, experiment templates, and variant histories.
• Manage clinical data, including PHI, in accordance with security, privacy, and compliance standards.
• Bachelor's degree in Data Science, Computer Science, Statistics, or a related quantitative discipline — or equivalent experience.
• Over 5 years of experience in data science or applied machine learning, including the deployment and maintenance of models or AI systems in production environments.
• At least 2 years of recent, hands-on experience in evaluating and enhancing LLM-based systems.
• Proficient in structured error analysis, prompt/configuration iteration, experiment design, and interpretation of metrics.
• Familiarity with LLM evaluation methodologies and tools, including golden/regression sets, LLM-as-judge with validation, and tracing/observability tools.
• Strong Python skills and solid data analysis capabilities; knowledge of SQL is an advantage.
• Comfortable computing and reasoning about sensitivity, specificity, and positive predictive value (PPV).
• Hypothesis-driven approach to work.
• Excellent written communication skills.
• Must manage clinical data, including PHI, in compliance with security and privacy regulations.
• Comprehensive background check is mandatory.
• In-person I-9 verification is required.
• May be subject to drug screening prior to employment.
• Government-issued photo ID is necessary for identity verification.
• High-speed internet connection of over 10 Mbps at home is required.
• Preferred qualifications include: healthcare experience; mentoring or technical leadership; experience collaborating with clinical reviewers or domain experts; familiarity with LLM observability/telemetry stacks; background in statistics or experimentation; Master's degree in a quantitative field.
• Enjoy a healthy work/life balance.
• Benefit from flexible work arrangements and autonomy.
• Access comprehensive benefits, including health insurance options, for eligible employees.
• High-speed home internet requirements are supported as part of the home-working capability.
Tendios
GuidePoint Security
Seneca Holdings
Get handpicked remote jobs straight to your inbox weekly.