
Principal AI Engineering Architect
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in United States.
• Establish the technical strategy and oversee architectural design for cloud, data, and AI/ML systems in end-to-end projects, taking ownership of architectural decisions and guiding solutions from research to large-scale production.
• Design and implement production-grade multi-agent AI systems, encompassing agent orchestration, tool usage, memory management, and inter-agent communication frameworks.
• Develop and implement solutions using Amazon Bedrock AgentCore and other complementary AWS GenAI services to securely deploy, scale, and operate agentic workloads in production.
• Create scalable cloud-native solutions with a strong preference for AWS, incorporating multi-cloud and hybrid strategies as necessary (AWS as primary, with Azure, GCP, and Kubernetes as secondary options).
• Design data architectures including data warehouses, data lakes, and pipelines for both batch and streaming workloads (e.g., Snowflake, Redshift, BigQuery, Spark, Kafka).
• Create AI/ML systems that include model serving, MLOps pipelines, feature stores, and LLM-based applications (e.g., SageMaker, Bedrock, AgentCore, Vertex AI, MLflow, Hugging Face).
• Build and enhance scalable ML platforms, pipelines, and infrastructure that facilitate reliable and repeatable model development and deployment across various teams.
• Define standards for infrastructure as code, CI/CD, and DevOps across projects (e.g., Terraform, CloudFormation, GitHub Actions).
• Drive optimization for performance, scalability, cost efficiency, and reliability across deployed systems.
• Ensure that the architecture complies with security, governance, and compliance standards (e.g., GDPR, HIPAA, SOC2).
• Lead initiatives for cloud migrations and platform modernization.
• Set benchmarks for AI-forward engineering, utilizing tools like Claude and Cursor effectively and aiding the team in their adoption.
• Over 8 years of software engineering experience, including at least 5 years in technical leadership positions and 4+ years focused on AI/ML systems in a production environment.
• Strong background in software engineering (Python or a similar language) with a keen design sense for creating scalable and maintainable systems.
• Extensive hands-on experience in designing and deploying production multi-agent AI systems, including agent orchestration, planning, tool usage, and multi-agent coordination.
• Profound expertise in AWS, with detailed knowledge of AWS GenAI offerings and practical experience with Amazon Bedrock AgentCore; experience with other cloud platforms (Azure, GCP) is advantageous.
• Solid background in microservices, serverless architecture, containers, and event-driven systems (e.g., Kubernetes, Docker, Lambda, EventBridge).
• Proficient in infrastructure as code and CI/CD methodologies (e.g., Terraform, CloudFormation, Pulumi, GitHub Actions).
• Strong data architecture knowledge across relational, NoSQL, and big data systems (e.g., PostgreSQL, MongoDB, Snowflake, BigQuery, Spark, Kafka).
• Practical experience with data modeling, ETL/ELT pipelines, and orchestration tools (e.g., Airflow, Prefect, dbt).
• Expertise in AI frameworks and orchestration tools for constructing agentic systems (e.g., LangChain, LangGraph, AgentCore, CrewAI, AutoGen, or similar tools).
• Significant experience in designing AI/ML systems for production, including LLMs, MLOps, and model serving (e.g., SageMaker, Bedrock, Vertex AI, MLflow, Hugging Face, PyTorch, TensorFlow).
• Strong background in evaluation frameworks and observability tools for LLM and agentic applications, including the development of such capabilities where they are lacking.
• Comprehensive understanding of AI safety, responsible AI principles, prompt injection defenses, and handling of PII.
• Extensive experience in building RAG pipelines: chunking strategies, embedding models, vector databases, and advanced retrieval methods.
• API design experience, including the architecture and integration of internal and external services at scale.
• Advanced expertise in cost optimization strategies: token economics, caching methods, model routing, and quantization.
• Solid grasp of networking, security, identity, and access management within cloud environments.
• Experience with governance, compliance, and observability frameworks.
• Proven track record of senior technical leadership and mentoring seasoned engineers.
• Excellent communication skills with stakeholders, capable of translating complex technical concepts to various audiences.
• Demonstrated daily usage and expert knowledge of AI-forward coding tools such as Claude Code and Cursor.
• Experience in multi-cloud architecture, AI ethics or responsible AI practices, or holding enterprise architecture certifications (e.g., TOGAF, AWS/Azure/GCP) is a plus.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible work arrangements and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative and innovative work environment.
Get handpicked remote jobs straight to your inbox weekly.