
Senior DevOps Engineer, EKS/Kubernetes
Posted Jul 22

Posted Jul 22
This is a fully remote position, open to applicants in United States.
• Design, implement, and sustain Linux infrastructure in both on-premises and cloud settings.
• Automate infrastructure setup and configuration using tools like Terraform, Ansible, or CloudFormation.
• Oversee and enhance AWS environments with a focus on performance, scalability, security, and cost-effectiveness.
• Establish and maintain monitoring, logging, and alerting solutions (e.g., Datadog, Prometheus, Grafana, ELK, CloudWatch).
• Architect, deploy, and manage production Kubernetes/AWS EKS clusters, encompassing node group strategy, cluster upgrades, multi-tenant workload isolation, and cross-region disaster recovery (DR) architecture and implementations.
• Define and spearhead cluster upgrade, security hardening, and disaster recovery initiatives for large-scale production Kubernetes/AWS EKS environments, while acting as a senior technical resource for intricate production incidents.
• Manage Kubernetes networking, including VPC CNI configuration and ingress controllers (ALB/NGINX/Traefik).
• Implement IAM roles for pod security standards and network policies to safeguard EKS workloads.
• Configure and optimize cluster autoscaling (Cluster Autoscaler or Karpenter) and workload autoscaling (HPA/VPA) to enhance performance and cost efficiency.
• Develop and sustain Helm charts and GitOps-based deployment pipelines (e.g., ArgoCD, Flux) for Kubernetes workloads.
• Manage Docker container builds and registries to support EKS-based application deployments.
• Deploy, scale, and maintain GitLab Runners (including Kubernetes executor runners on EKS) to enhance CI/CD pipeline throughput and reliability.
• Support and assist in managing database platforms on AWS RDS (MySQL, PostgreSQL), working with data owners to ensure performance and reliability.
• Ensure that systems adhere to security and compliance standards, including SOX and SOC 2 initiatives.
• Execute and maintain Linux patching strategies, addressing security updates and CVEs promptly.
• Engage in incident response, root cause analysis, and recovery operations.
• Collaborate with development, QA, and cross-functional teams to enhance reliability, release processes, and operational standards.
• Participate in on-call rotations and offer after-hours support as necessary.
• Bachelor’s degree in Computer Science, Information Technology, or a related discipline.
• 8+ years of experience in Linux Systems Administration, DevOps, or Site Reliability Engineering roles.
• 5+ years of experience with AWS services, including EC2, VPC, IAM, RDS, S3, and CloudWatch.
• 5+ years of hands-on experience in designing and managing production workloads on Kubernetes/AWS EKS, covering cluster upgrades, networking, and autoscaling.
• Proficiency in scripting and automation using Python and Bash.
• Extensive hands-on experience with Infrastructure as Code using Terraform, and with CI/CD pipelines (GitLab CI/CD), including executing CI/CD workloads on Kubernetes/EKS.
• Proficiency in Docker, Helm, and Kubernetes troubleshooting in a production environment.
• Strong understanding of networking fundamentals and cloud security best practices.
• CKA (Certified Kubernetes Administrator) certification is expected or actively pursued; Preferred qualifications include CKAD and AWS certifications (e.g., AWS Certified DevOps Engineer, Solutions Architect) are advantageous.
• Experience with Karpenter, Kyverno, OPA/Gatekeeper, Falco, and multi-cluster/multi-tenant EKS environments.
• Familiarity with AI and automation tools such as Claude, Cursor, and OpenAI to enhance DevOps workflows through AI-assisted CI/CD, self-healing operations, and automated incident response.
• Experience with microservices, serverless architectures, and DevSecOps practices.
• Highly competitive and inclusive medical, dental, and vision coverage options.
• Health Savings Account for medical and dependent care expenses.
• Flexible Spending Account to cover specific out-of-pocket expenses.
• Paid time off, including vacation, sick leave, and holidays.
• 401k match and financial planning tools.
• LTD and STD insurance coverages, along with optional benefits.
• Employee Assistance Program.
• Pet Insurance.
• Legal Assistance.
• Tuition Assistance.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.