
Director, AI Evaluation
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Pennsylvania.
β’ The Director of AI Evaluation is responsible for defining, demonstrating, and maintaining quality across the entire AI portfolio at Geisinger, encompassing both internally developed models and vendor-supplied systems.
β’ Oversees the team involved in building and validating AI systems: this includes data scientists who create production-level machine learning and refined AI systems, as well as senior analysts who assess their performance.
β’ Establishes quality standards, leads the team in enforcing these standards, and communicates findings to the VP of AI, executive leadership, and the board.
β’ Responsible for the methodology that ensures initiatives meet these standards, including pre-production validation and ongoing production monitoring.
β’ Attracts and retains top technical talent, fosters career development, and leads the evaluation team.
β’ Offers practical technical guidance to program teams while they design validation studies, conduct equity audits, develop monitoring plans, and create escalation playbooks.
β’ Manages the evaluation toolkit and reusable playbooks and templates, enabling each new program to progress more swiftly than its predecessor.
β’ Links each AI initiative to the outcomes it was intended to enhance, using a pre-launch baseline for reference.
β’ Experience in people leadership, including managing, developing, and cultivating technical staff; focuses on building cohesive teams rather than merely leading projects.
β’ Strong background in experimental design and causal inference, with the ability to discern which methods are appropriate for various scenarios.
β’ Practical experience in designing and executing model evaluation studies in real-world production environments.
β’ Familiarity with evaluating large language models (LLM) or generative AI systems, or equivalent experience with complex machine learning systems characterized by ambiguous or noisy ground truth.
β’ Demonstrated capability to convert unclear failure modes into specific, defensible evaluation frameworks and monitoring metrics.
β’ Proficient in Python and SQL, with a working knowledge of contemporary ML tools and cloud-native data environments.
β’ Experience in assessing fairness and equity within machine learning systems.
β’ Excellent written communication skills, as the role involves producing evaluation memos and specifications that are relied upon by non-technical decision-makers.
β’ Experience in healthcare, clinical settings, or regulated industries is highly preferred.
β’ We provide healthcare benefits for both full-time and part-time roles starting from day one, which includes vision, dental, and domestic partner coverage.
β’ We promote a collaborative, cooperative, and collegial work environment.
Anyone AI
EWOR
Jerry
Get handpicked remote jobs straight to your inbox weekly.