Lead DevOps Engineer

atICFRemoteUS flagVirginiaFull-timeDevOps & Site Reliability Engineer (SRE)Senior$131.3k – $223.1k/year

Posted 1 day ago

This is a fully remote position, open to applicants in Virginia.

📋 Description

• Take ownership of the architecture and daily management of multi-account AWS environments across development, QA, staging, and production.

• Create and maintain reusable Terraform modules and Terragrunt environment setups.

• Manage Amazon EKS Auto Mode and Kubernetes resources utilizing Helm, Argo CD, and ApplicationSets.

• Provide support for GitOps controllers and automate Kubernetes infrastructure.

• Develop and uphold GitHub Actions workflows, AWS OIDC authentication, self-hosted runners, container-image pipelines, environment promotion, deployment approvals, and rollback strategies.

• Operate Microsoft SQL Server on Amazon RDS, including tasks like backups, recovery, monitoring, and AWS DMS migration workflows.

• Enforce security and audit controls using AWS security services and centralized logging.

• Establish and maintain observability through Prometheus, Grafana, CloudWatch, OpenTelemetry, dashboards, metrics, logs, traces, alerts, and service-level objectives.

• Collaborate with developers to review and troubleshoot Java and Spring Boot services.

• Incorporate application and supply-chain security into delivery pipelines.

• Direct incident response, root-cause analysis, disaster recovery drills, infrastructure upgrades, cost optimization, architecture documentation, operational runbooks, and knowledge transfer.


⛳️ Requirements

• A Bachelor’s degree in computer science, information technology, engineering, or a related discipline, or equivalent professional experience.

• Over 8 years of experience in DevOps, cloud infrastructure, platform engineering, or site reliability.

• At least 5 years of experience managing production workloads on AWS.

• A minimum of 5 years of experience with Terraform and infrastructure as code (IaC).

• Over 4 years of experience administering production Kubernetes environments, including Amazon EKS, Helm, ingress, RBAC, secrets, upgrades, observability, and GitOps delivery with Argo CD or a similar platform.

• At least 3 years of experience building and supporting CI/CD pipelines using GitHub Actions or a comparable platform.

• Strong skills in Linux, scripting, AWS networking, IAM, troubleshooting, and incident response.

• Must be a US Citizen or Permanent Resident in accordance with contract requirements.

• Familiarity with Terragrunt, EKS Auto Mode, AWS Pod Identity, Argo CD ApplicationSets, AWS Controllers for Kubernetes, or Kubernetes Resource Orchestrator.

• Experience managing Microsoft SQL Server on Amazon RDS and supporting AWS DMS, backup and restore functionalities, disaster recovery, and database performance monitoring.

• Knowledge of CloudFront, Route 53, ACM, Cognito, SQS, SNS, SES, Secrets Manager, S3, ECR, and cross-account AWS access patterns.

• Experience in implementing observability with Prometheus, Grafana, CloudWatch, OpenTelemetry, centralized logging, alerting, and service-level objectives.

• Experience with SAST using SonarQube, DAST, software composition analysis, container scanning, SBOMs, and policy-based release gates.

• Hands-on experience in reviewing and troubleshooting Java and Spring Boot applications, including REST or GraphQL APIs, Maven, JVM diagnostics, database connection pools, and containerized deployments.


🏝️ Benefits

• Equal opportunity employer.

• Reasonable accommodations for disabled veterans, individuals with disabilities, and individuals with sincerely held religious beliefs.

• Confidential accommodation support.

• Benefit offerings referenced through Transparency in Coverage Act.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers