
Principal DevOps Architect
Posted Sep 16

Posted Sep 16
This is a fully remote position, open to applicants in United States.
• Take ownership of the platform architecture and technical roadmap concerning infrastructure, deployment, observability, and the AI/ML platform.
• Establish engineering standards, design patterns, and optimal pathways for Infrastructure as Code (IaC), Continuous Integration/Continuous Deployment (CI/CD), and AI tooling.
• Foster adoption through reference implementations and architecture evaluations.
• Manage cloud infrastructure as code using Terraform, which includes reusable modules, remote state management, peer-reviewed pull requests, drift detection, and automated plan/apply processes in CI/CD.
• Implement policy-as-code for security and cost management guardrails.
• Design and deploy AWS infrastructure across development, UAT, staging, and production environments.
• Construct and maintain CI/CD pipelines for large-scale AWS applications.
• Oversee release management, including rollback strategies, blue/green deployments, canary deployments, and release gates.
• Package and execute containerized workloads utilizing Docker and Kubernetes (EKS).
• Lead the SLI/SLO/SLA program along with observability initiatives using OpenTelemetry.
• Minimize Mean Time To Detect (MTTD) and Mean Time To Recover (MTTR), conduct blameless post-incident reviews, and participate in on-call duties.
• Provision and manage the AI/ML platform utilizing Anthropic Claude via AWS Bedrock and internal MCP services.
• Develop LLMOps practices encompassing prompt versioning, evaluation pipelines, token cost attribution, guardrails, and agent-action audit logging.
• Enforce the protection of PHI data boundaries to BAA-covered providers.
• Operate and demonstrate controls for SOC 2 Type 2 and HIPAA compliance.
• Manage secrets, ensure supply-chain security, maintain Software Bill of Materials (SBOM) and conduct image/dependency scanning, along with FinOps responsibilities.
• Bachelor's degree in Software Engineering or a related field, or a comparable combination of technical education and professional experience.
• Over 10 years of experience in Site Reliability Engineering (SRE), DevOps, or Platform Engineering, including time spent at a senior individual contributor or architect level (Staff, Principal, or Architect).
• Established technical authority across teams without direct management responsibilities.
• Proven track record of driving the adoption of new practices or platforms across multiple teams.
• Proficient in Terraform, including the use of reusable modules, remote state management, and CI/CD change management.
• Experience in building and operating CI/CD pipelines for large-scale applications hosted on AWS.
• Familiarity with GitHub Actions, Jenkins, GitLab, or AWS-native CI/CD solutions.
• Experience managing containerized workloads using Docker and Kubernetes.
• Skills in monitoring and troubleshooting with cloud-native tools and OpenTelemetry.
• Competence in Linux system administration, Unix scripting, and automation techniques.
• Experience in environments that require HIPAA, HITECH, HITRUST, PHI, PII, or PCI DSS compliance.
• Remote work flexibility.
• Engaging hands-on individual-contributor role.
• Participation in on-call rotation.
• Opportunity for professional influence through architecture and technical expertise.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.