
Staff DevOps Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Architect and continuously enhance Azure DevOps CI/CD pipelines, incorporating pipeline-as-code standards, templating strategies, and artifact promotion workflows.
• Take ownership of the health and evolution of AWS EKS clusters, managing node lifecycle, autoscaling, networking, RBAC, and cluster upgrades.
• Design and implement Infrastructure-as-Code practices while advocating for GitOps patterns.
• Propel platform reliability enhancements utilizing observability data from New Relic in collaboration with SRE.
• Define and maintain standardized templates for containerized workloads, including Dockerfile standards and Helm chart libraries.
• Collaborate with engineering teams to onboard new services and minimize manual work through automation.
• Act as an escalation point for complex infrastructure incidents via PagerDuty and partake in the on-call rotation.
• Lead reviews following incidents and implement systemic fixes to decrease page volume and MTTR.
• Maintain runbooks and document platform processes in Confluence.
• Define and promote DevOps standards throughout the engineering organization.
• Conduct architecture reviews and provide technical advice on infrastructure decisions.
• Mentor senior and mid-level engineers through pairing, code reviews, and knowledge sharing.
• Identify gaps in tooling and develop business cases for platform investments.
• 8+ years of experience in DevOps or platform engineering.
• A minimum of 2 years in a Staff or Principal role within an organization of over 100 engineers.
• Extensive, hands-on experience with Kubernetes; EKS experience is particularly preferred.
• Experience in troubleshooting workloads, networking, storage, and cluster operations at scale.
• Strong proficiency in Azure DevOps Pipelines, including YAML pipeline authoring, library management, service connections, and environment promotion gates.
• Proven track record in designing and maintaining CI/CD systems for microservice architectures with multiple independent teams.
• Experience with operating observability platforms such as New Relic or Datadog to facilitate proactive reliability improvements.
• Proficiency in at least one scripting language: Python, Bash, or Go.
• Expertise with Infrastructure-as-Code tools: Terraform, Pulumi, or CDK.
• Familiarity with feature flag patterns and progressive delivery; experience with Unleash or similar tools is a plus.
• Exceptional written communication skills, with the ability to translate complex infrastructure decisions into actionable guidance.
• Preferred: experience with data engineering or ML infrastructure workloads on Kubernetes.
• Preferred: background in contributing to or maintaining internal developer portals such as Backstage.
• Preferred: familiarity with FinOps practices and AWS cost attribution and optimization.
• Preferred: experience in SRE-related roles and comfort with SLO/SLI definitions and error budget policies.
• Comprehensive medical coverage.
• Dental coverage.
• Vision coverage.
• 401(k) match.
• Generous PTO.
• Paid parental leave.
• Health and wellness programs.
• Tuition reimbursement.
• Employee discounts.
OnePay
Arista Networks
Octus
Tandem Diabetes Care
Get handpicked remote jobs straight to your inbox weekly.