
Staff DevOps Engineer
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in California.
• Act as a technical cornerstone for the platform engineering discipline.
• Design and enhance CI/CD pipelines within Azure DevOps, including standards for pipeline-as-code, templating strategies, and workflows for artifact promotion.
• Oversee the health and development of AWS EKS clusters, focusing on node lifecycle management, autoscaling, networking, RBAC, and cluster upgrades.
• Create and enforce Infrastructure-as-Code methodologies while promoting GitOps practices.
• Propel platform reliability enhancements using observability data from New Relic in collaboration with SRE.
• Define and maintain standard templates for containerized workloads, encompassing Dockerfile benchmarks and Helm chart libraries.
• Collaborate with engineering teams to onboard services and minimize manual processes through automation.
• Serve as an escalation point for complex infrastructure incidents through PagerDuty.
• Participate in the on-call rotation and lead post-incident analysis.
• Identify recurring failure patterns and implement systemic solutions to decrease page volume and mean time to recovery (MTTR).
• Maintain runbooks and documentation for the platform in Confluence.
• Establish and promote DevOps standards throughout the engineering organization.
• Conduct architecture evaluations and offer technical guidance.
• Mentor senior and mid-level engineers through pairing, code reviews, and knowledge sharing.
• Identify gaps in tooling and create business cases for platform investments.
• 8+ years of experience in DevOps or platform engineering.
• A minimum of 2 years working at a Staff or Principal level within an organization of over 100 engineers.
• Extensive hands-on expertise with Kubernetes, including troubleshooting workloads, networking, storage, and cluster operations at scale.
• Experience with AWS EKS is preferred.
• Strong proficiency in Azure DevOps Pipelines, including YAML pipeline creation, library management, service connections, and environment promotion gates.
• Demonstrated experience in designing and maintaining CI/CD systems for microservice architectures involving multiple independent teams.
• Experience with observability platforms such as New Relic or Datadog.
• Proficiency in programming languages like Python, Bash, or Go.
• Familiarity with Infrastructure-as-Code tools such as Terraform, Pulumi, or CDK.
• Understanding of feature flag methodologies and progressive delivery; experience with Unleash or similar tools is a plus.
• Excellent written communication abilities.
• Preferred: experience with data engineering or machine learning infrastructure workloads on Kubernetes.
• Preferred: background in contributing to or maintaining internal developer portals like Backstage.
• Preferred: knowledge of FinOps practices and AWS cost attribution and optimization.
• Preferred: experience in roles adjacent to SRE, with comfort in defining SLO/SLI and managing error budget policies.
• Comprehensive medical insurance.
• Dental insurance.
• Vision insurance.
• 401(k) matching program.
• Generous paid time off (PTO).
• Paid parental leave.
• Health and wellness initiatives.
• Tuition reimbursement program.
• Employee discounts.
Truelogic Software
Cadwell
Raya
Arize AI
Get handpicked remote jobs straight to your inbox weekly.