Staff Infrastructure Engineer

Posted Aug 13

This is a fully remote position, open to applicants in California, +2 more states.

📋 Description

• Design and take ownership of the cloud platform utilized by all Headway engineers.

• Revamp the deployment architecture to prevent failures and limit their impact.

• Manage the ECS and EKS container usage while assessing wider EKS adoption for AI tasks.

• Create the next version of inter-service network connectivity.

• Oversee capacity engineering for variable workloads and enhance autoscaling dependability.

• Develop a Terraform self-service infrastructure platform with established guardrails.

• Implement cost attribution per team across AWS, Datadog, and LLM expenditures.

• Manage the performance of the Python monolith runtime, focusing on garbage collection, event loop contention, and runtime constraints.

• Lead the upgrades of frameworks and packages.

• Offer technical guidance through architecture evaluations, runbooks, and standardized tooling.

• Influence infrastructure choices across teams and elevate the overall infrastructure engineering standards within the organization.


⛳️ Requirements

• A minimum of 8 years in platform, infrastructure, or SRE roles at organizations handling substantial production traffic.

• Extensive AWS knowledge and hands-on experience with compute and networking at scale, including ECS, EKS, RDS, networking, and IAM.

• Proficient in infrastructure-as-code, particularly with Terraform.

• Background in designing self-service infrastructure platforms for engineering teams.

• Practical experience in autoscaling and capacity engineering.

• Familiarity with container orchestration using ECS and/or EKS.

• Proven history of making deployments safe and self-service for other teams.

• Ability to influence at the staff level across team boundaries without needing management authority.

• Experience in FinOps and cloud cost optimization is advantageous.

• In-depth knowledge of Kubernetes and EKS is a plus.

• Experience with scaling observability tools, such as Datadog, is beneficial.

• Background in healthcare or regulated environments is a plus.

• Familiarity with event-driven systems is a nice-to-have.

• Authorization to work in the United States is required.

• Willingness or capability to address employment visa sponsorship needs.


🏝️ Benefits

• Comprehensive and competitive total rewards package.

• Extensive health and wellness benefits.

• Retirement savings options.

• Significant ownership opportunities through equity.

• Reasonable accommodations for individuals with disabilities.

People also viewed

adconova GmbH1 day ago

Senior Cloud & AI Infrastructure Engineer

DE flagGermany OnlyFull-timeInfrastructure Engineer€60k – €100k/year
ApplyView job
Pragmatike2 days ago

IoT & Cloud Infrastructure Engineer

KR flagSouth Korea OnlyFull-timeInfrastructure Engineer
ApplyView job
Pragmatike2 days ago

IoT & Cloud Infrastructure Engineer

CN flagChina OnlyFull-timeInfrastructure Engineer
ApplyView job
Aalyria3 days ago

Senior Software Engineer – Infrastructure & Platform

US flagCalifornia OnlyFull-timeInfrastructure Engineer$165k – $215k/year
ApplyView job
appsoluts GmbH3 days ago

Web, Backend & Infrastructure Developer

DE flagGermany OnlyFull-timeInfrastructure Engineer€50k – €75k/year
ApplyView job
SYNCREON3 days ago

Senior Mainframe Systems Programmer – Infrastructure Engineering

US flagNew Jersey OnlyFull-timeInfrastructure Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers