
Staff AI Architect
Posted 21 hours ago

Posted 21 hours ago
This is a fully remote position, open to applicants in United States.
• Define the architecture for AI systems and lead the design across applications involving LLM, agentic systems, retrieval, and ML model serving for comprehensive engagements.
• Drive the strategy for RAG and context-engineering, which includes hybrid retrieval, reranking, chunking, and the selection of embedding models across vector stores.
• Lead the design of multi-agent orchestration utilizing frameworks such as LangGraph, Amazon Bedrock Agents, or AgentCore.
• Oversee the model serving and inference architecture on AWS, enhancing latency, throughput, and cost through streaming, caching, batching, and quantization.
• Establish LLMOps and MLOps standards across all engagements, including prompt versioning, evaluation pipelines, CI/CD, and feature stores.
• Drive the evaluation, guardrails, and observability strategy, creating both offline and online evaluation methods, LLM-as-judge scoring, and tracing mechanisms.
• Manage the responsible AI framework, which encompasses governance, PII and data isolation, prompt-injection defenses, and grounding controls.
• Make decisions regarding fine-tuning, prompting, RAG, and tool usage while considering accuracy, cost, and maintainability.
• Utilize AI-driven tools like Claude and Cursor to deliver high-quality work quickly.
• Collaborate with leadership and clients on the technical direction of AI, translating business objectives into architectural decisions.
• Convey intricate AI concepts and tradeoffs to both engineering and non-engineering stakeholders.
• Work in partnership with engineering, data, ML, and product teams.
• Create and maintain documentation on AI architecture, evaluation standards, and reusable patterns.
• Set architectural standards and best practices within the team.
• Provide mentorship to junior and mid-level engineers.
• Serve as a technical escalation point for complex architectural and integration challenges.
• Assess emerging technologies and recommend appropriate tools, frameworks, and patterns.
• Over 7 years of professional software engineering experience.
• A minimum of 3 years in AI/ML architecture or technical leadership roles.
• Extensive expertise in LLM and agentic application architecture, including multi-agent orchestration and tool-calling patterns.
• Strong experience in designing production RAG and retrieval systems, including hybrid retrieval, reranking, and selecting vector stores.
• Significant experience with model serving and inference optimization on AWS, including tuning for latency, throughput, and cost.
• In-depth understanding of LLMOps and MLOps, encompassing evaluation pipelines, prompt versioning, CI/CD, and observability for AI workloads.
• Proven experience with AI evaluation, guardrails, and tracing tools.
• Solid judgment regarding the trade-offs among fine-tuning, prompting, and RAG, focusing on accuracy, cost, and maintainability.
• Expert software engineering background, with proficiency in Python and at least one additional language (e.g., TypeScript, Java, Go).
• Strong foundation in cloud-native architecture, including microservices, serverless computing, containerization, and event-driven systems.
• Comprehensive understanding of responsible AI and governance, including data isolation, prompt-injection defense, and compliance measures.
• Demonstrated leadership and mentoring experience within a team context.
• Regular, hands-on use and expert knowledge of AI-driven tools such as Claude and Cursor.
• Exceptional problem-solving abilities and capacity to navigate complex technical and business challenges.
• AWS AI certifications, experience in fine-tuning or customizing models, or expertise in responsible AI are advantageous.
• Successful completion of a background check may be necessary.
• Paid time off.
• Medical insurance.
• Dental insurance.
• Vision insurance.
• 401(k).
Bullhorn
SAIC
Capital One
Capital One
Get handpicked remote jobs straight to your inbox weekly.