Principal DevOps Architect

Posted 23 hours ago

This is a fully remote position, open to applicants in Texas.

📋 Description

• Take ownership of the cloud platform architecture and technical roadmap for infrastructure, deployment, observability, and AI platforms.

• Provide guidance to the VP of Software Architecture and engineering leadership.

• Establish engineering standards, patterns, golden paths, and reference implementations for infrastructure as code, CI/CD, and AI tooling.

• Serve as a hands-on technical authority, mentoring engineers across teams without direct management responsibilities.

• Promote the adoption of infrastructure as code and AI tooling while preventing the unregulated use of unverified AI tools.

• Advise on build/buy decisions and assess platform-tooling vendors.

• Manage AWS infrastructure as code using Terraform, including reusable modules, remote state, drift detection, and automated plan/apply processes.

• Implement policy-as-code security and cost management measures.

• Design and deploy AWS infrastructure for development, UAT, staging, and production environments.

• Maintain a multi-tenant database and hosting architecture, ensuring replication and per-tenant isolation.

• Construct and operate CI/CD pipelines with automated build, testing, and deployment processes.

• Oversee release management, rollback strategies, deployment gates, and customer acceptance procedures.

• Manage Docker and Kubernetes/EKS workloads, automating configurations with Ansible and Packer.

• Lead SLI/SLO/SLA initiatives, manage error budgets, and utilize OpenTelemetry for observability.

• Analyze production events, conduct blameless post-incident reviews, participate in on-call rotations, and resolve incidents.

• Provision and manage AI/ML platforms using Anthropic Claude via AWS Bedrock or the Anthropic API.

• Establish AI workload guardrails, observability, cost controls, usage auditing, and ensure compliance with PHI boundaries.

• Operate agentic and MCP tooling with least-privilege access and a human-in-the-loop escalation process.

• Develop LLMOps practices, including prompt versioning, evaluation frameworks, token-cost attribution, and audit logging.

• Route inference across multiple model providers and manage protocols such as MCP.

• Maintain and demonstrate SOC 2 Type 2 and HIPAA controls and assist with external audits.

• Handle secrets management, supply-chain security, vulnerability remediation, SBOMs, cloud costs, tagging, budgets, and right-sizing.

• Collaborate with product, development, support, operations, and QA teams across a distributed, multi-time-zone environment.


⛳️ Requirements

• Bachelor’s degree in software engineering or a relevant combination of technical education and professional experience.

• Over 10 years of experience in SRE/DevOps/platform engineering, including senior individual-contributor or architect-level roles.

• Proven experience in delivering CI/CD, REST API deployment, containerization, IaaS/PaaS, data pipelines, and application observability.

• Demonstrated technical authority across teams, influencing through expertise and example rather than direct management.

• Experience in driving the adoption of infrastructure as code, CI/CD transformations, or AI/ML platforms across multiple teams.

• Proficiency in Terraform, including reusable modules, remote state management, and CI/CD change management.

• Experience in building and maintaining CI/CD pipelines for large-scale AWS applications using GitHub Actions, Jenkins, GitLab, or AWS-native tools.

• Experience in managing Docker and Kubernetes containerized workloads.

• Familiarity with cloud-native monitoring, troubleshooting, and OpenTelemetry.

• Knowledge of Linux system administration, Unix scripting, and automation techniques.

• Experience in environments governed by HIPAA, HITECH, HITRUST, PHI, PII, or PCI DSS regulations.

• At least 2 years of experience in PHP, MySQL, and SQL is advantageous.

• Experience with AI/ML or GenAI production platforms is a plus.

• Familiarity with LLMOps, MCP services, AI security/governance, policy-as-code, secrets management, supply-chain security, FinOps, clinical research, or healthcare technology is beneficial.

• Candidates must successfully complete and pass reference and background checks.

• Applicants must indicate their desired salary for consideration.

• Must verify identity and eligibility to work in compliance with federal regulations.


🏝️ Benefits

• Health insurance.

• Long-term disability insurance.

• Life insurance.

• Unlimited Paid Time Off.

• 10 paid holidays.

• Paid parental leave.

• Work anniversary bonus.

• Participation in the Employee of the Quarter Program.

• Monthly $100 connectivity stipend reimbursement.

• 401(k) matching: 100% of the first 3% invested and 50% of the next 2%.

• Remote and telecommuting work arrangement.

People also viewed

knowmad mood15 hours ago

Consultor/a DevSecOps – AWS

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Koniag Government Services1 day ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Koniag Government Services1 day ago

Senior AWS DevOps Engineer – AWS, Kubernetes, HCP, CI/CD, Observability, AI-focus

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ASRC Federal1 day ago

Senior DevOps Administrator – Supporting NASA

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nelios1 day ago

DevOps Engineer, Cloud Infrastructure

GR flagGreece OnlyPart-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Worth AI1 day ago

Senior DevOps Engineer, Infrastructure – Reliability

US flagFlorida OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers