
AI Product Lead – AI Experience, Quality
Posted 7 hours ago

Posted 7 hours ago
This is a fully remote position, open to applicants in California.
• Oversee AI behavior that interacts with members throughout the processes of prompt creation, agentic editing, segmentation, scheduling, analytics, and recommendations.
• Develop, test, version, and document prompts and agent instructions.
• Specify the necessary context, member data, tools, and product state, translating these into requirements and acceptance criteria.
• Manage the AI evaluation process, including datasets, behavioral scenarios, scoring rubrics, regression coverage, and human review.
• Conduct regular evaluations and quality monitoring; establish baselines and release gates for changes in prompts, models, and agents.
• Analyze traces and AI interactions to identify failures related to prompts, context, models, tools, product structure, and implementation.
• Resolve identified issues and offer actionable recommendations for engineering teams.
• Transform member research, support feedback, analytics, production failures, and domain expertise into prioritized tasks, experiments, and regression cases.
• Direct the daily product strategy and learning agenda for the AI experience.
• Collaborate with product, design, marketing, copy, and engineering teams.
• Demonstrated experience in shaping or significantly enhancing real LLM or agentic products, showing proof of improved user outcomes or product quality.
• Practical understanding of prompts, context, tools, agent behavior, and evaluation within production experiences.
• Direct experience with prompt iteration, evaluation tools, and trace analysis.
• Proficient in working with JSON, schemas, and structured outputs.
• Strong product judgment, with the ability to translate ambiguous member needs into clear behaviors, quality definitions, and testable hypotheses.
• Commitment to experimental rigor, including the use of representative test sets and failure taxonomies.
• Systems thinking with the capability to create reusable patterns.
• Excellent communication skills across both technical and non-technical domains.
• Ability to make trade-offs involving member value, quality, latency, reliability, and cost considerations.
• While not required to build or deploy production code, must be able to operate the prompt and evaluation system independently.
• Bonus: experience with AI product behavior, conversation design, computational linguistics, or prompt systems.
• Bonus: experience in developing content-generation, creative, or marketing tools.
• Bonus: experience with Langfuse, Braintrust, or similar evaluation and observability platforms.
• Bonus: experience working on products targeted at small businesses, creators, or marketers.
• Comprehensive health insurance coverage for individuals.
• 16 weeks of paid parental leave for non-birthing parents.
• 22 weeks of paid maternity leave for birthing parents.
• Unlimited flexible time off.
• 401(k) matching (applicable to US employees only).
• $1,000 annual stipend for professional development and learning opportunities.
DIRECTV
Scale Army Careers
Braiins
Get handpicked remote jobs straight to your inbox weekly.