Remotery

Site Reliability Engineer – US Central/Eastern Time

Posted Aug 6

This is a fully remote position, open to applicants in United States.

📋 Description

• Transform a rapidly growing, stateful system into a reliable, fully automated platform by managing provisioning, scaling, rebalancing, and recovery.

• Manage EKS clusters across multiple environments utilizing Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments.

• Oversee and develop a multi-AWS-account organization, including tasks related to provisioning, networking, access control, and cross-account connectivity.

• Sustain the Terraform/Terragrunt infrastructure-as-code framework, which encompasses modules, plan-on-PR/apply-on-merge pipelines, and secure shared-infrastructure patterns.

• Enhance operational tools for deployments, schema modifications, backups, restorations, and incident management.

• Recognize recurrent operational challenges and resolve them through coding and self-healing automation.

• Streamline cloud expenditure.

• Engage in on-call duties and incident management, while progressively decreasing incident occurrences.

• Design and automate the platform layer that underpins the organization’s services.


⛳️ Requirements

• Extensive hands-on experience with Kubernetes in a production environment; EKS is preferred.

• Proficient in diagnosing node pressure, networking challenges, and deployment failures at scale, including environments with thousands of nodes.

• Strong background in managing production infrastructure on AWS across various accounts.

• Knowledge of AWS organizational boundaries, IAM, and inter-account networking.

• Experience in automating infrastructure using Terraform or Terragrunt at scale, including module design and state management.

• Solid understanding of Linux systems, including disk, memory, networking, and failure modes.

• Experience supporting stateful systems such as databases, queues, and storage solutions.

• Ability to troubleshoot and analyze production performance and reliability challenges.

• Comfortable taking ownership of systems end-to-end, including on-call duties.

• Located in the US Central or Eastern timezone.


🏝️ Benefits

• Flexible remote work arrangement.

• Meeting-free days on Tuesdays and Thursdays.

• Prioritization of focused building time over perfect coordination.

• Default async communication.

• Fair and accessible interview process with available accommodations or adjustments.

People also viewed

CVS Health10 hours ago

Salesforce DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$83.4k – $166.9k/year
ApplyView job
Devoteam11 hours ago

Data, AWS DevSecOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aspirion12 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Goodgame Studios12 hours ago

Senior Agentic Engineer – Java Backend, DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Instacart12 hours ago

Site Reliability Engineer II

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$133k – $169k/year
ApplyView job
Logicalis Spain13 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€40k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers