
Staff AI Platform Engineer – Agent & Retrieval Infrastructure
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Develop the integration for Amazon Bedrock, encompassing agent and action group setup, backend APIs, model access, throughput, and cross-environment deployment.
• Take ownership of the retrieval pipeline from data ingestion and chunking to embedding and storage within Amazon OpenSearch Serverless.
• Enhance index design, cost efficiency, and capacity.
• Modify ingestion pipelines for internal knowledge, oceanic data, and customer platforms.
• Tackle challenges related to geospatial and large binary datasets.
• Establish security measures with Bedrock Guardrails, VPC and PrivateLink boundaries, least-privilege IAM, and audit trails.
• Specify governance protocols for approval boundaries and secure autonomous agent operations.
• Execute LLMOps and monitoring for tracing, tool invocations, and retrieval performance using CloudWatch and tools like Langfuse or Phoenix.
• Create an automated evaluation infrastructure, monitor outcomes, and manage model-accuracy release thresholds.
• Provide abstraction layers, SDKs, and self-service environments for autonomous AI feature deployment.
• Oversee infrastructure as code across various environments.
• Guarantee deployment safety and engage in incident reviews.
• Define the sequence of internal, operational, ocean-data, and customer-facing retrieval functionalities.
• Collaborate with the core data transport team to establish retrieval and data freshness requirements.
• Over 8 years of experience in software and infrastructure engineering.
• Extensive production backend experience; proficiency in Python or TypeScript is preferred, while Go is acceptable.
• Staff-level responsibility for technical direction.
• Hands-on experience deploying Amazon Bedrock in production, including agents, knowledge bases, guardrails, model access, throughput, and quotas.
• Experience in containerized service deployment on ECS, EKS, or Lambda.
• Ownership of CI/CD processes.
• Practical experience with RAG and vector search, including embeddings, chunking strategies, semantic search quality, and managed vector databases at scale and cost.
• Experience in building or significantly enhancing ingestion pipelines for messy, heterogeneous, and unstructured data sources.
• Strong expertise in AWS services: IAM, VPC networking, PrivateLink, Lambda, S3, KMS, CloudWatch, and infrastructure as code using Terraform, CDK, or CloudFormation.
• Production experience with LLM features or autonomous agents.
• Experience securing agentic systems, including managing tool permissions, prompt injection risks, data exfiltration, sensitive data handling, and human-in-the-loop controls.
• Experience designing developer-facing APIs, SDKs, or platform services with an API-first approach.
• Experience building and maintaining multi-tenant services with customer data isolation.
• Proven staff-level technical leadership and system design expertise.
• Pragmatic preference for managed infrastructure solutions.
• Ability to collaborate across multiple functions in a small team and document systems comprehensively.
• Legally authorized to work in the United States.
• Eligible for obtaining a government security clearance if necessary.
• Equity.
• Equal opportunity employer.
EverCommerce
ARETUM
Civica US
Vannevar Labs
Get handpicked remote jobs straight to your inbox weekly.