
Staff Applied AI Engineer – Product & Agent Performance
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in United States.
• Create and refine agent behaviors within actual, live workflows, encompassing long-horizon, multi-turn agentic tasks.
• Develop retrieval and context architecture to ensure agents remain anchored in real data.
• Design memory and state management for multi-turn and multi-agent interactions.
• Construct context and prompt templates utilizing few-shot examples, structured formatting, and reasoning scaffolding.
• Enhance performance through effective prompting, strategic tool usage, and context development.
• Conduct evaluations under real production conditions.
• Draft evaluation rubrics, quality heuristics, and severity- and cost-weighted thresholds.
• Create and validate escalation pathways for human review based on confidence and uncertainty levels.
• Optimize performance with a focus on cost-awareness, latency, reliability, and accuracy.
• Assess and approve model modifications, making go/no-go decisions for launches.
• Maintain comprehensive AI documentation at the product level, including model cards, intended use, limitations, and known failure modes.
• Collaborate with Product teams and product managers to ensure agents are steerable, trustworthy, and prepared for scaling.
• Establish baselines grounded in production, document failure modes, and develop measurement strategies.
• Analyze retrieval, context, memory, and escalation patterns to identify areas for reliability, calibration, and cost enhancements.
• Transform production insights into actionable recommendations for Product and Engineering.
• Equivalent practical experience demonstrating the depth required for this senior-level position.
• Over 8 years of experience in production software engineering.
• More than 3 years of direct experience managing ML, LLM, or agentic systems in a production environment.
• Proven experience in healthcare, finance, or another regulated industry.
• Expertise in diagnosing agent failures and linking solutions to instruction, retrieval, context, or memory design.
• Capability to assess failures by severity and cost rather than merely by frequency.
• Practical experience with RAG architecture.
• Familiarity with production-focused evaluation frameworks.
• Experience implementing fallback or human-in-the-loop logic in automated systems.
• Working knowledge of AWS AI/ML services, including Bedrock and SageMaker.
• Experience applying AI in healthcare settings or workflows where safety, transparency, and calibrated uncertainty impact care teams or patients (preferred).
• Experience with long-horizon, multi-turn or multi-agent workflows and product-level AI documentation such as model cards (preferred).
• The chance to shape how agent performance, safety, and readiness are evaluated for production healthcare workflows.
• Significant ownership over prompts, context, memory, evaluations, and escalation patterns at scale.
• A cross-functional role that translates production evidence into AI enhancements utilized across Arcadia’s platform.
• Work for a mission-driven company dedicated to improving patient care.
• Enjoy a flexible, remote-friendly culture that values personality and heart.
• Participate in employee-driven programs and initiatives aimed at personal and professional growth.
• Become a member of the talented, energized, diverse, and purpose-driven Arcadian community.
Get handpicked remote jobs straight to your inbox weekly.