Remotery

Senior DevOps Engineer, EKS/Kubernetes

Posted Jul 22

This is a fully remote position, open to applicants in United States.

📋 Description

• Design, implement, and sustain Linux infrastructure in both on-premises and cloud settings.

• Automate infrastructure setup and configuration using tools like Terraform, Ansible, or CloudFormation.

• Oversee and enhance AWS environments with a focus on performance, scalability, security, and cost-effectiveness.

• Establish and maintain monitoring, logging, and alerting solutions (e.g., Datadog, Prometheus, Grafana, ELK, CloudWatch).

• Architect, deploy, and manage production Kubernetes/AWS EKS clusters, encompassing node group strategy, cluster upgrades, multi-tenant workload isolation, and cross-region disaster recovery (DR) architecture and implementations.

• Define and spearhead cluster upgrade, security hardening, and disaster recovery initiatives for large-scale production Kubernetes/AWS EKS environments, while acting as a senior technical resource for intricate production incidents.

• Manage Kubernetes networking, including VPC CNI configuration and ingress controllers (ALB/NGINX/Traefik).

• Implement IAM roles for pod security standards and network policies to safeguard EKS workloads.

• Configure and optimize cluster autoscaling (Cluster Autoscaler or Karpenter) and workload autoscaling (HPA/VPA) to enhance performance and cost efficiency.

• Develop and sustain Helm charts and GitOps-based deployment pipelines (e.g., ArgoCD, Flux) for Kubernetes workloads.

• Manage Docker container builds and registries to support EKS-based application deployments.

• Deploy, scale, and maintain GitLab Runners (including Kubernetes executor runners on EKS) to enhance CI/CD pipeline throughput and reliability.

• Support and assist in managing database platforms on AWS RDS (MySQL, PostgreSQL), working with data owners to ensure performance and reliability.

• Ensure that systems adhere to security and compliance standards, including SOX and SOC 2 initiatives.

• Execute and maintain Linux patching strategies, addressing security updates and CVEs promptly.

• Engage in incident response, root cause analysis, and recovery operations.

• Collaborate with development, QA, and cross-functional teams to enhance reliability, release processes, and operational standards.

• Participate in on-call rotations and offer after-hours support as necessary.


⛳️ Requirements

• Bachelor’s degree in Computer Science, Information Technology, or a related discipline.

• 8+ years of experience in Linux Systems Administration, DevOps, or Site Reliability Engineering roles.

• 5+ years of experience with AWS services, including EC2, VPC, IAM, RDS, S3, and CloudWatch.

• 5+ years of hands-on experience in designing and managing production workloads on Kubernetes/AWS EKS, covering cluster upgrades, networking, and autoscaling.

• Proficiency in scripting and automation using Python and Bash.

• Extensive hands-on experience with Infrastructure as Code using Terraform, and with CI/CD pipelines (GitLab CI/CD), including executing CI/CD workloads on Kubernetes/EKS.

• Proficiency in Docker, Helm, and Kubernetes troubleshooting in a production environment.

• Strong understanding of networking fundamentals and cloud security best practices.

• CKA (Certified Kubernetes Administrator) certification is expected or actively pursued; Preferred qualifications include CKAD and AWS certifications (e.g., AWS Certified DevOps Engineer, Solutions Architect) are advantageous.

• Experience with Karpenter, Kyverno, OPA/Gatekeeper, Falco, and multi-cluster/multi-tenant EKS environments.

• Familiarity with AI and automation tools such as Claude, Cursor, and OpenAI to enhance DevOps workflows through AI-assisted CI/CD, self-healing operations, and automated incident response.

• Experience with microservices, serverless architectures, and DevSecOps practices.


🏝️ Benefits

• Highly competitive and inclusive medical, dental, and vision coverage options.

• Health Savings Account for medical and dependent care expenses.

• Flexible Spending Account to cover specific out-of-pocket expenses.

• Paid time off, including vacation, sick leave, and holidays.

• 401k match and financial planning tools.

• LTD and STD insurance coverages, along with optional benefits.

• Employee Assistance Program.

• Pet Insurance.

• Legal Assistance.

• Tuition Assistance.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers