
Staff DevOps Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Design and systematically enhance Azure DevOps CI/CD pipelines, encompassing pipeline-as-code standards, templating strategies, and artifact promotion workflows.
• Take ownership of the health and development of AWS EKS clusters, focusing on node lifecycle management, autoscaling, networking, RBAC, and upgrades.
• Create and enforce Infrastructure-as-Code methodologies while promoting GitOps patterns.
• Propel platform reliability enhancements using New Relic observability data in collaboration with SRE.
• Establish and maintain golden-path templates for containerized workloads, including Dockerfile standards and Helm chart libraries.
• Collaborate with engineering teams to onboard services and mitigate manual work through automation.
• Manage and coordinate complex infrastructure incidents via PagerDuty, take part in on-call rotations, and lead post-incident evaluations.
• Detect recurring failure modes and implement solutions that minimize page volume and MTTR.
• Keep runbooks and platform documentation updated in Confluence.
• Define and promote DevOps standards regarding pipelines, containers, secrets, and deployment safety.
• Conduct architecture evaluations and offer technical advice.
• Guide senior and mid-level engineers through pairing, code reviews, and knowledge transfer.
• Recognize tooling deficiencies and develop business cases for platform investments.
• Over 8 years of experience in DevOps or platform engineering.
• Minimum of 2 years working at a Staff or Principal level within an organization comprising 100+ engineers.
• Extensive, hands-on knowledge of Kubernetes, particularly with AWS EKS, covering workloads, networking, storage, and cluster operations at scale.
• Strong expertise in Azure DevOps Pipelines, including YAML pipeline creation, library management, service connections, and environment promotion gates.
• Demonstrated experience in designing and maintaining CI/CD systems for microservice architectures with multiple independent teams.
• Experience managing observability platforms such as New Relic or Datadog to facilitate proactive reliability enhancements.
• Proficiency in programming languages like Python, Bash, or Go.
• Skilled with Infrastructure-as-Code tools such as Terraform, Pulumi, or CDK.
• Familiarity with feature flagging patterns and progressive delivery; experience with Unleash or similar tools is a plus.
• Exceptional written communication skills.
• Experience with data engineering or ML infrastructure workloads on Kubernetes is preferred.
• Background in contributing to or managing internal developer portals is preferred.
• Understanding of FinOps practices and tools for AWS cost attribution and optimization is preferred.
• Experience in SRE-adjacent roles and comfort with SLO/SLI definitions and error budget policies is preferred.
• Comprehensive medical, dental, and vision coverage.
• 401(k) matching.
• Generous paid time off (PTO).
• Paid parental leave.
• Health and wellness initiatives.
• Tuition reimbursement.
• Employee discounts.
OnePay
Arista Networks
Octus
Tandem Diabetes Care
Get handpicked remote jobs straight to your inbox weekly.