
Software Engineer, AI Systems
Posted Sep 28

Posted Sep 28
This is a fully remote position, open to applicants in United States.
• Construct and manage multi-step LLM pipelines that coordinate model calls, tool calls, graph queries, retrieval processes, quality gates, and specialist-agent transitions.
• Enhance Haven’s coordinated agent team and orchestration layer from incident evidence through analysis, review, and enterprise learning.
• Develop the context layer to facilitate Neo4j graph traversal, vector search, and hybrid retrieval.
• Create evaluation datasets, scoring systems, regression suites, model comparisons, human-label loops, and quality attribution for each stage.
• Execute tracing, tool-call audits, cost and latency monitoring, failure management, and quality dashboards.
• Identify loops, hallucinations, and silent drift prior to customer detection.
• Choose models from OpenAI, Anthropic, and Google based on specific task needs.
• Collaborate with the product and knowledge engineering teams.
• Contribute to shaping the AI roadmap.
• Design, deploy, instrument, and refine AI/LLM reasoning systems using production evidence.
• Report directly to the CTO.
• Experience in AI or ML engineering, including the deployment of LLM systems relied upon by real users.
• Proficient in building and troubleshooting multi-step, tool-calling workflows using LangGraph, LangChain, or a comparable framework.
• Adopt a repeatable methodology for LLM evaluation, encompassing representative datasets, regression testing, LLM-as-judge techniques, or human review loops.
• Expertise in assembling context for LLMs with a clear perspective on what to retrieve, the volume, and the rationale behind it.
• Proven track record of managing systems from deployment through monitoring and incident response, including diagnosing and rectifying failures or regressions.
• Comfortable navigating various model providers and articulating trade-offs in quality, latency, cost, context, and operational risk.
• Familiarity with Neo4j and Cypher, or a similar graph store, with a strong ability to quickly learn graph data modeling (highly preferred).
• Proficient in Python.
• Production experience with FastAPI, asynchronous services, testing, observability, and maintainable interfaces.
• Familiarity with Cypher fluency, schema evolution, MERGE patterns, embeddings, and operating a live knowledge graph (nice to have).
• Awareness of prompt-injection, context-leak prevention, tenant isolation, role-based access, and policy-layer separation (nice to have).
• Experience in running production AI services on Azure and utilizing Azure AI Search, Pinecone, MongoDB Atlas, pgvector, Elasticsearch, or a similar platform (nice to have).
• Previous experience in an enterprise-level environment (strongly preferred).
• Candidates must be based in the US; relocation is not an option.
• Engage with meaningful reasoning challenges.
• Experience an evaluation-first culture.
• See visible impact on customers.
• Work within a small team with high ownership.
• Receive direct feedback from safety teams and users.
• Collaborate closely with the CTO, product, and knowledge engineering teams.
• Have the opportunity to make significant technical decisions and witness your work reach customers swiftly.
Brillio
Gainwell Technologies
SmartLogic
Get handpicked remote jobs straight to your inbox weekly.