
AI Systems Architect
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in United States.
• Take ownership of the core infrastructure: Establish the technical strategy for media pipelines, LLM routing, distributed data systems, and observability tools that support the entire platform, designing them from foundational principles.
• Create AI agent frameworks: Develop AI agents and multi-agent systems from the ground up, including orchestration, tool invocation, MCP servers, A2A dispatch, and the necessary permissions and observability infrastructure to ensure reliability in production.
• Delve into distributed systems: Tackle the challenging aspects such as data flow, fault boundaries, consistency trade-offs, and performance under load to construct resilient infrastructure that performs when it is most needed.
• Enhance reliability standards: Incorporate observability from the outset, utilizing tracing, metrics, and structured logging, while developing quality systems that enable a fast-paced team to operate without disruptions.
• Influence the roadmap: Provide an engineering perspective to platform strategy, determining what to build, the order of development, and the rationale behind these decisions.
• Have established technical direction across various teams, with the ability to highlight specific systems you have designed and operated in production, along with instances where you successfully guided others.
• Possess deep expertise in distributed systems, including the underlying layers of the framework — data flow, fault tolerance, consistency, and performance under real-world conditions.
• Design AI agents and agentic systems: Orchestration, tool invocation, multi-agent coordination, and the infrastructure that ensures observability and production readiness.
• Have constructed large-scale platforms from the ground up, making foundational decisions and living with their consequences, rather than merely joining an established codebase.
• Develop clean, robust APIs across REST, gRPC, event-driven, and schema-based systems, and are proficient with MCP, A2A, schema registries, and tool-calling normalization.
• Have a thorough understanding of media pipelines — including transcoding, ffmpeg-class tools, cloud storage, and large-object processing — with strong hands-on experience in PostgreSQL, OpenSearch, Redis, Docker, Kubernetes, and Terraform or Pulumi.
• Be the go-to person that teams prefer in discussions when difficult architectural decisions need to be made.
• Competitive compensation
• Comprehensive benefits
• Generous access to frontier and open-weight models for daily tasks, prototyping, and evaluations
OpenText
Akamai Technologies
Professional Physical Therapy
Get handpicked remote jobs straight to your inbox weekly.