
Lead Kubernetes Platform Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Oversee a two-month EKS modernization discovery initiative
• Establish a baseline for the multi-cluster production EKS environment, which includes inventory, topology, workload placement, ownership, cost, Kubernetes versions, and support status
• Evaluate failure domains, isolation boundaries, and dependency concentrations
• Create a multi-cluster target architecture along with workload placement and tenancy models
• Design and execute automated EKS upgrades, utilizing blue/green or in-place methods, Karpenter drift-based node rotation, add-on compatibility, and API deprecation management
• Lead instance sizing, workload-fit analysis, capacity planning, bin-packing, network-limit analysis, and scale-up headroom planning
• Strategize and manage phased ARM64/Graviton migration for Java, Go, PHP, and Python workloads
• Set up Karpenter NodePools, including weights, overlays, and effective Spot capacity patterns
• Model compute, support, and commitment economics while quantifying potential savings
• Evaluate and enhance ingress, Gateway API, service mesh, CNI, GitOps, and infrastructure-as-code strategies
• Measure maintenance and toil, establish reduction targets, and develop automation solutions
• Define an AI-assisted DevOps agent workstream and assess platform readiness
• Lead the TechPod as a player-coach, acting as the primary liaison for customer infrastructure leadership
• Collaborate with AWS specialists on capacity planning and architectural decisions
• Create estate baselines, architecture designs, upgrade plans, roadmaps, and executive summaries
• Present findings and recommendations to engineering leadership
• 8+ years of experience in DevOps, SRE, Platform, or Infrastructure Engineering
• 4+ years of hands-on experience operating production Kubernetes environments
• Previous experience in a technical lead, staff, or principal role
• Extensive production experience with Amazon EKS at a large scale
• Direct ownership of EKS version upgrades across multiple production clusters
• Advanced experience with Karpenter (v1+)
• Strong understanding of EC2 instance families, generations, CPU architectures, and network performance limitations
• Experience in migrating production workloads to Graviton or other ARM64 platforms
• Proficient in modeling compute costs and commitments, including Savings Plans, Reserved Instances, and Spot
• Solid knowledge of AWS VPC CNI, ingress controllers, Gateway API, service mesh, and EKS load balancing
• Advanced skills in Terraform
• Experience with Helm and GitOps tools like Argo CD or Flux
• Comfortable using Datadog, Prometheus, Grafana, or similar tools
• Strong scripting capabilities in Python, Go, or Bash
• Ability to produce a current-state overview and roadmap in a matter of weeks
• Ability to translate technical findings into cost, risk, and capacity terms
• U.S. work authorization and sponsorship-status questions will be required in the application process
• Fully Remote Workplace
• Unlimited Paid Time Off
• Equity Opportunities
• 401K with company contributions
• Sponsored healthcare coverage
• Training and certification programs for professional development
T-Rex Solutions, LLC
Takeaway.com
Dadoteca
Abercrombie
Get handpicked remote jobs straight to your inbox weekly.