
Director, AI Automation Engineering
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in New York.
• Oversee the infrastructure that consistently validates whether autonomous agents perform intended actions prior to and following production.
• Stage, test, and manage the lifecycle of bots and agent workflows.
• Establish baseline quality metrics and coverage standards for AI projects.
• Develop infrastructure to execute automated research chains and agent evaluations on a large scale.
• Create automated evaluation frameworks for LLM outputs that assess accuracy, safety, bias, and regression.
• Implement RAG evaluation, conduct prompt regression testing, and establish golden-set benchmarking.
• Define and uphold quality thresholds prior to release.
• Collaborate with AI Researchers and Engineers on evaluation design for new model and agent patterns.
• Construct release-gating infrastructure, which includes staged rollouts and rollback authority.
• Develop AI automation solutions to ensure software quality across the product portfolio.
• Integrate testing into GitHub Actions or similar platforms for AI model and agent releases.
• Maintain test suites as models, prompts, and tools evolve.
• Work alongside AI Solutions Architects, AI Engineers, Principal AI Product, and product teams on evaluation, launch readiness, QA, and feature validation.
• Document standards, runbooks, and patterns to enable non-engineers using Sage to conduct basic testing independently.
• Set technical standards and progress towards leading a small automation team.
• A minimum of 5 years of experience in building, evaluating, or governing production AI/ML systems, with direct exposure to agentic AI, LLM-based automation, or autonomous system safety.
• Hands-on experience evaluating LLM outputs, including RAG evaluation, prompt regression, safety, and accuracy benchmarking.
• Proven experience in constructing CI/CD pipelines with integrated test gates using GitHub Actions or similar tools.
• Comfortable working across AI platform infrastructure and product-oriented AI features.
• Capability to work independently as a senior individual contributor and convey quality standards to both engineers and non-engineers.
• Experience in agent testing, bot lifecycle management, or AI agent orchestration platforms is highly desirable.
• Familiarity with LLM observability and evaluation tools such as LangSmith, Braintrust, or custom harnesses is a significant advantage.
• Background in healthcare, pharmaceuticals, or regulated environments is a strong plus.
• Previous experience in establishing a QA or automation function from the ground up, with a desire to evolve into a leadership role.
• Proficiency in Python and standard testing frameworks like pytest or equivalent.
• Medical, dental, and vision insurance coverage for you and your dependents.
• Access to an on-demand healthcare concierge.
• Pre-tax savings options for HSA, FSA, and DCFSA.
• Monthly employer contributions to HSA if enrolled in a high-deductible health plan.
• Fully paid short- and long-term disability insurance.
• Life and AD&D insurance coverage.
• Flexible vacation policy.
• Paid parental leave after 6 months of employment.
• Option to work remotely.
• 401(k) plan with company matching.
Element Fleet Management
Leidos
The Amatriot Group
Foresite Cybersecurity
Get handpicked remote jobs straight to your inbox weekly.