
Senior Applied AI/ML Scientist
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Take charge of the complete development process for agents and agentic features, encompassing requirements, success criteria, evaluations, guardrails, and tools.
• Direct the selection, development, and validation of models.
• Create and conduct A/B tests and other meticulous experiments to evaluate customer impact.
• Establish and refine the quality standards for agents.
• Collaborate with engineers who are building and maintaining agents and evaluation infrastructure.
• Recognize opportunities for enhancing agent-platform capabilities, including prompt enhancements, reusable judges, context-window management, tool utilization, and adoption of cutting-edge models.
• Articulate AI capabilities, limitations, and quality to product, design, engineering, and leadership teams.
• Be responsible for the evaluation and quality of agents and lead evidence-based decisions for model selection and improvement.
• Contribute to the comprehensive AI quality and evaluation strategy and promote best practices in evaluation.
• Foster relationships with business unit leaders, stakeholders, and executives.
• Promote technical excellence and remain updated with the latest research advancements.
• Over 5 years of practical AI/ML experience with a proven history in evaluation, measurement, and quality of implemented AI/ML or LLM systems.
• Experience in designing evaluation frameworks and developing reliable judges or automated evaluators for ML or LLM/agentic systems.
• Strong discernment in model selection and validation.
• Comprehensive experimentation and statistical expertise, including A/B testing, power analysis, and effective measurement design.
• Foundational knowledge of ML and modeling, including proficiency in Python, common ML frameworks, transformers, and embeddings.
• Proficient in SQL and comfortable working with large datasets.
• Capability to collaborate closely with product, design, and engineering teams, translating technical quality queries into actionable insights.
• Preferred direct experience in evaluating LLM-driven or agentic products in a production environment.
• Preferred experience in defining quality standards or evaluation methodologies.
• Preferred skills in rapid prototyping with new AI/ML techniques.
• Familiarity with evaluation tooling, observability, and MLOps practices is preferred.
• Experience in influencing product roadmaps with data regarding AI quality and feasibility is preferred.
• Awareness of model security, bias mitigation, and responsible AI practices is preferred.
• Premium BCBSIL medical, dental, and vision insurance for you and eligible dependents.
• Full, free access to Modern Health, including coaching, therapy sessions, and digital wellness resources.
• 401(k) plan with a 50% company match on the first 6% of contributions.
• 100% employer-paid Life and Disability insurance.
• Flexible PTO policy along with company-wide Rest & Recharge days.
• Up to 16 weeks of paid parental leave.
• $1,000 USD annual Lifestyle Spending Account.
• One-time $550 USD home office setup stipend.
• Monthly $50 USD internet stipend.
• 16 hours of paid volunteer time each year.
• $100 annual charitable donation match.
• Pre-tax commuter benefits.
• Subsidized child/eldercare services through Care.com.
• Discounted pet insurance via Figo.
• Complimentary personalized financial wellness support through Your Money Line.
NVIDIA
Jazz Pharmaceuticals
Hydra Host
Grupo Consciência
Get handpicked remote jobs straight to your inbox weekly.