
AI Agent Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Pakistan.
β’ Develop and enhance production AI agents utilizing foundation models, specifically AWS Bedrock.
β’ Create and refine orchestration for production AI agents.
β’ Implement context engineering within production LLM and agent systems.
β’ Construct and manage evaluation datasets and pipelines, including tool selection, trajectory assessments, and LLM-as-judge evaluations.
β’ Establish and uphold production observability and monitoring for LLM and agent systems.
β’ Execute tracing and instrumentation for production LLM systems.
β’ Collaborate directly with the engineer currently spearheading the AI-agenting function as the inaugural dedicated hire in this domain.
β’ Utilize agent-architecture principles to determine if the existing custom orchestration layer should be preserved or if a production framework should be adopted.
β’ Participate as a member of a new three-person product team alongside the AI function lead and a Full Stack Engineer.
β’ Function as an individual contributor rather than overseeing others.
β’ Assume responsibility for measurable deliverables right from the onboarding process.
β’ Deliver comparable AI/agent engineering tasks in alignment with roadmap deadlines.
β’ Over 2 years of experience in developing production LLM agents, encompassing tool-calling agent loops, streaming, context management, structured outputs, and orchestration frameworks.
β’ Practical experience with an agentic framework such as LangGraph, LangChain, or custom orchestration.
β’ At least 1.5 years of experience in evaluation-driven development, which includes building and managing evaluation datasets and pipelines for tool selection, trajectory evaluation, and LLM-as-judge.
β’ A minimum of 1 year of experience with LLM observability, tracing, and instrumentation utilizing Langfuse, OpenTelemetry, or similar tools.
β’ Authentic production agent-observability experience.
β’ Direct, hands-on experience with Langfuse is highly preferred; OpenTelemetry or other tracing tools are acceptable only as a secondary indicator alongside genuine agent-observability exposure.
β’ Over 1 year of experience in LLM cost optimization, incorporating prompt caching, model selection and routing, as well as LLM FinOps.
β’ More than 5 years of backend engineering expertise, including proficiency in TypeScript/Node, Postgres, and serverless AWS.
β’ Demonstrated success in delivering similar projects within comparable timelines.
β’ Experience in deploying agentic AI systems in production, with the capability to discuss specific failure modes and their mitigations while operating end-to-end across the AI stack.
β’ Fixed Shifts: 11:30 AM - 9:00 PM PKT (Summer) | 12:30 PM - 10:00 PM PKT (Winter)
β’ No Weekend Work: True work-life balance, not just a phrase.
β’ Day 1 Benefits: Laptop and other assets provided.
β’ Support That Matters: Mentorship, community, and platforms for idea sharing.
β’ True Belonging: A long-term career where your contributions are acknowledged.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.