Lead DevOps Engineer

atICFRemoteUS flagVirginiaFull-timeDevOps & Site Reliability Engineer (SRE)Senior$131.3k – $223.1k/year

Posted 13 hours ago

This is a fully remote position, open to applicants in Virginia.

📋 Description

• Take charge of the architecture and daily management of multi-account AWS environments spanning development, QA, staging, and production.

• Create and sustain reusable Terraform modules along with Terragrunt environment configurations.

• Manage Amazon EKS Auto Mode and Kubernetes resources utilizing Helm, Argo CD, and ApplicationSets.

• Assist with GitOps controllers and automate Kubernetes infrastructure.

• Design and uphold GitHub Actions workflows, AWS OIDC authentication, self-hosted runners, container-image pipelines, environment promotion, deployment approvals, and rollback procedures.

• Oversee Microsoft SQL Server on Amazon RDS and facilitate AWS DMS migration workflows.

• Enforce security and audit measures using AWS security services and centralized logging.

• Establish and maintain observability using Prometheus, Grafana, CloudWatch, OpenTelemetry, dashboards, metrics, logs, traces, alerts, and service-level objectives.

• Collaborate with developers to assess and resolve issues in Java and Spring Boot services.

• Embed application and supply-chain security within delivery pipelines.

• Spearhead incident response, root-cause analysis, disaster recovery exercises, infrastructure upgrades, cost optimization, architecture documentation, operational runbooks, and knowledge transfer.


⛳️ Requirements

• A Bachelor's degree in computer science, information technology, engineering, or a related field, or equivalent professional experience.

• Over 8 years of experience in DevOps, cloud infrastructure, platform engineering, or site reliability, with at least 5 years managing production workloads on AWS.

• More than 5 years of hands-on experience with Terraform and infrastructure as code (IaC), covering reusable modules, remote state, imports, drift reconciliation, and automated plan and apply workflows.

• At least 4 years of experience managing production Kubernetes environments, including Amazon EKS, Helm, ingress, RBAC, secrets, upgrades, observability, and GitOps delivery using Argo CD or a similar platform.

• A minimum of 3 years of experience in developing and supporting CI/CD pipelines with GitHub Actions or a comparable platform.

• Proficient in Linux, scripting, AWS networking, IAM, troubleshooting, and incident response.

• Must be a US Citizen or Permanent Resident as per contract requirements.

• Familiarity with Terragrunt, EKS Auto Mode, AWS Pod Identity, Argo CD ApplicationSets, AWS Controllers for Kubernetes, or Kubernetes Resource Orchestrator.

• Experience in managing Microsoft SQL Server on Amazon RDS and supporting AWS DMS, backup and restore, disaster recovery, and database performance monitoring.

• Knowledge of CloudFront, Route 53, ACM, Cognito, SQS, SNS, SES, Secrets Manager, S3, ECR, and cross-account AWS access patterns.

• Expertise in implementing observability with Prometheus, Grafana, CloudWatch, OpenTelemetry, centralized logging, alerting, and service-level objectives.

• Experience applying SAST with SonarQube, DAST, software composition analysis, container scanning, SBOMs, and policy-based release gates.

• Practical experience in reviewing and troubleshooting Java and Spring Boot applications, including REST or GraphQL APIs, Maven, JVM diagnostics, database connection pools, and containerized deployments.


🏝️ Benefits

• Equal opportunity employer.

• Reasonable accommodations for disabled veterans, individuals with disabilities, and individuals with sincerely held religious beliefs.

• Confidential accommodation assistance.

• Benefit offerings covered under the Transparency in (Benefits) Coverage Act.

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Mastercam17 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers