
Site Reliability Engineer – US Central/Eastern Time
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in United States.
• Transform a rapidly growing, stateful system into a reliable, fully automated platform by managing provisioning, scaling, rebalancing, and recovery.
• Manage EKS clusters across multiple environments utilizing Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments.
• Oversee and develop a multi-AWS-account organization, including tasks related to provisioning, networking, access control, and cross-account connectivity.
• Sustain the Terraform/Terragrunt infrastructure-as-code framework, which encompasses modules, plan-on-PR/apply-on-merge pipelines, and secure shared-infrastructure patterns.
• Enhance operational tools for deployments, schema modifications, backups, restorations, and incident management.
• Recognize recurrent operational challenges and resolve them through coding and self-healing automation.
• Streamline cloud expenditure.
• Engage in on-call duties and incident management, while progressively decreasing incident occurrences.
• Design and automate the platform layer that underpins the organization’s services.
• Extensive hands-on experience with Kubernetes in a production environment; EKS is preferred.
• Proficient in diagnosing node pressure, networking challenges, and deployment failures at scale, including environments with thousands of nodes.
• Strong background in managing production infrastructure on AWS across various accounts.
• Knowledge of AWS organizational boundaries, IAM, and inter-account networking.
• Experience in automating infrastructure using Terraform or Terragrunt at scale, including module design and state management.
• Solid understanding of Linux systems, including disk, memory, networking, and failure modes.
• Experience supporting stateful systems such as databases, queues, and storage solutions.
• Ability to troubleshoot and analyze production performance and reliability challenges.
• Comfortable taking ownership of systems end-to-end, including on-call duties.
• Located in the US Central or Eastern timezone.
• Flexible remote work arrangement.
• Meeting-free days on Tuesdays and Thursdays.
• Prioritization of focused building time over perfect coordination.
• Default async communication.
• Fair and accessible interview process with available accommodations or adjustments.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.