
Technical Product Manager, Data Science
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in California.
• Take responsibility for the precision, dependability, and cost-effectiveness of the AI supporting CentrlX.
• Develop and manage the evaluation framework, which includes creating golden datasets, grading rubrics, and LLM-as-judge pipelines aligned with human assessments, as well as regression suites.
• Establish quantifiable quality benchmarks for document digitization, extraction, retrieval, groundedness, summaries, responses, evaluations, and multi-step agent workflows.
• Assess agent performance, focusing on tool selection, retrieval accuracy, step sequencing, and the quality of final deliverables.
• Transform client setbacks into permanent evaluation cases.
• Conduct benchmarking of models across various providers regarding accuracy, latency, and cost.
• Oversee model migrations and make decisions on model deployment based on workflow requirements.
• Monitor AI expenditures and enhance cost efficiency through strategies like routing, model tiering, caching, and context optimization.
• Specify logging and tracing requirements for prompts, retrieved contexts, outputs, tool calls, token counts, and latency metrics.
• Collaborate with Product and Design teams to implement in-app feedback capture.
• Create, label, curate, and sustain evaluation datasets with holdouts and rotations.
• Act as the primary contact for AI quality escalations; manage triage, reproduction, root cause analysis, and resolution.
• Compose user stories and acceptance criteria while executing changes to prompts, configurations, and models.
• Regularly publish quality and cost reports.
• Collaborate with domain experts to incorporate industry insights into evaluation rubrics.
• Must possess work authorization in the USA.
• Minimum of 3 years of experience in product management, data science, or applied AI, including at least 2 years focused on LLM-based products in a production environment.
• Proficient in Python and SQL.
• Comfortable utilizing a notebook for data extraction, batch inference, and metric computation.
• Proven experience in constructing evaluation datasets and harnesses for LLM systems, including golden sets, rubric development, LLM-as-judge with human calibration, and regression testing in response to prompt and model adjustments.
• Working proficiency with at least one evaluation or LLM observability platform such as Braintrust, LangSmith, Langfuse, Arize Phoenix, W&B Weave, Inspect, Promptfoo, or a similar in-house solution.
• Practical knowledge of RAG systems, including retrieval quality, groundedness and faithfulness, hallucination detection, and chunking and context strategies.
• Statistical literacy to assess comparisons, evaluate significance, and recognize when differences are not substantive.
• Capability to articulate clear user stories and acceptance criteria while functioning within an agile engineering framework.
• Excellent written communication skills.
• Preferred: experience in document AI, agentic systems, financial services or investment management, inference cost reduction, in-product feedback systems, agent frameworks, and MCP.
• A degree in a quantitative or technical discipline is preferred.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Generous vacation and paid time off policies.
• Opportunities for professional development and continuous learning.
• Flexible work arrangements to promote work-life balance.
Emergent Software
DIRECTV
GE Vernova
DEUNA
Get handpicked remote jobs straight to your inbox weekly.