Staff DevOps Engineer

Posted 17 hours ago

This is a fully remote position, open to applicants in California.

📋 Description

• Design and continuously enhance Azure DevOps CI/CD pipelines, implementing pipeline-as-code standards, templating strategies, and workflows for artifact promotion.

• Oversee the health and progression of AWS EKS clusters, managing node lifecycle, autoscaling, networking, RBAC, and cluster upgrades.

• Create and enforce Infrastructure-as-Code methodologies utilizing Terraform or similar tools.

• Advocate for GitOps methodologies within engineering teams.

• Propel platform reliability enhancements by leveraging observability data from New Relic.

• Establish and maintain golden-path templates for containerized workloads, including standards for Dockerfiles and libraries for Helm charts.

• Collaborate with engineering teams to expedite service onboarding and minimize manual tasks through automation.

• Act as an escalation point for intricate infrastructure incidents managed through PagerDuty.

• Engage in on-call rotations and lead post-incident analyses for failures at the platform layer.

• Detect recurring failure patterns and implement systemic solutions to decrease page volume and Mean Time to Recovery (MTTR).

• Update and enhance runbooks and platform documentation within Confluence.

• Establish and communicate DevOps standards throughout the engineering organization.

• Conduct architecture reviews and offer technical guidance on decisions impacting infrastructure.

• Mentor senior and mid-level engineers through collaborative work, code review, and knowledge sharing.

• Identify gaps in tooling and develop business cases for platform investments.


⛳️ Requirements

• Over 8 years of experience in DevOps or platform engineering.

• Minimum of 2 years operating at a Staff or Principal level in an organization with 100+ engineers.

• Extensive, hands-on expertise in Kubernetes, preferably with EKS.

• Experience in troubleshooting Kubernetes workloads, networking, storage, and cluster operations at scale.

• Strong proficiency in Azure DevOps Pipelines, including YAML pipeline creation, library management, service connections, and environment promotion gates.

• Demonstrated experience in designing and maintaining CI/CD systems for microservice architectures involving multiple independent teams.

• Familiarity with operating observability platforms such as New Relic, Datadog, or similar solutions.

• Proficient in programming languages such as Python, Bash, or Go.

• Skilled in Infrastructure-as-Code tools like Terraform, Pulumi, or CDK.

• Awareness of feature flag patterns and progressive delivery; experience with Unleash or similar tools is a plus.

• Excellent written communication abilities.

• Experience in data engineering or machine learning infrastructure workloads on Kubernetes is preferred.

• Familiarity with internal developer portals like Backstage is preferred.

• Understanding of FinOps practices along with AWS cost attribution and optimization tools is preferred.

• Experience in SRE-adjacent roles with comfort in defining SLO/SLI and error budget policies is preferred.


🏝️ Benefits

• Comprehensive medical coverage.

• Dental coverage.

• Vision coverage.

• 401(k) matching.

• Generous paid time off (PTO).

• Paid parental leave.

• Health and wellness programs.

• Tuition reimbursement.

• Employee discounts.

People also viewed

Truelogic Software10 hours ago

Senior DevOps Engineer – Wealth Management Fintech

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Cadwell12 hours ago

Cloud Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $130k/year
ApplyView job
Raya13 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Arize AI14 hours ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$150k – $185k/year
ApplyView job
Capgemini16 hours ago

DevOps Engineer

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Cast & Crew17 hours ago

Staff DevOps Engineer

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$190k – $235k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers