
Applied AI Engineer
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in India.
• Take responsibility for the production improvement cycle, encompassing agent behavior, customer and operator feedback, evaluations, experiments, and validated business outcomes.
• Analyze agent workflows to gain insights into model interactions, tool usage, decision-making, failures, human modifications, and subsequent results.
• Establish quality benchmarks, representative evaluation datasets, regression coverage, and production monitoring processes.
• Explore the causes of agent underperformance, including context, knowledge, instructions, tools, routing, guardrails, and workflow design.
• Create and implement behavior enhancements that include prompting, context development, decision logic, tool utilization, and human review processes.
• Develop backend services, APIs, data models, and feedback systems to ensure observable, steerable, and reproducible agent behavior.
• Conduct controlled experiments, production replays, and phased rollouts.
• Collaborate with Product, Data Science, and Sales teams to identify high-value issues and define success for customers and the business.
• Implement safeguards for privacy, security, reliability, human oversight, and safe operational deployment.
• 2-5 years of experience in software engineering.
• Practical experience in building or managing LLM-powered features or agents in production environments, beyond just prototypes or demos.
• Proficiency in prompting and context engineering.
• Experience in developing or maintaining AI evaluation frameworks, including offline evaluation datasets, LLM-as-judge or human-in-the-loop scoring, and regression detection.
• Capability to analyze non-deterministic agent behavior in production traffic.
• Experience in executing A/B experiments, staged rollouts, or production replays.
• Proven history of delivering features that real users rely on and managing outcomes post-launch.
• High-agency mindset when investigating ambiguous performance issues across context, tools, routing, and workflow design.
• Familiarity with AI coding tools such as Cursor, Claude Code, Copilot, or similar.
• Preferred experience with voice or real-time conversational AI systems.
• Familiarity with LLM observability/tracing tools like Braintrust, LangSmith, or Datadog LLM Observability is preferred.
• Preferred experience with agentic orchestration frameworks such as LangChain/LangGraph or similar.
• Exposure to MCP-based tooling or agentic data workflows is preferred.
• A robust and competitive compensation package that includes a bonus and equity program.
• Flexible paid time off (PTO).
• 15 company holidays.
• 12 weeks of paid parental leave.
• 401k matching.
• An education stipend to support personal and professional development.
• Remote work stipend.
• A progressive benefits package for employees and their dependents.
• An open and transparent company culture.
futureproof consulting
Jones Lang LaSalle Americas, Inc.
UserTesting
Wearin'
Get handpicked remote jobs straight to your inbox weekly.