Staff DevOps Engineer

Posted Aug 27

This is a fully remote position, open to applicants in California.

📋 Description

• Take ownership of and enhance the fundamental infrastructure encompassing compute, networking, storage, and deployment systems.

• Design and manage CI/CD pipelines for AI, data, and product engineering teams.

• Develop and sustain infrastructure-as-code for consistent and auditable cloud, on-premises, and edge environments.

• Architect and oversee Kubernetes platforms tailored for training, inference, and application workloads, including GPU scheduling and autoscaling capabilities.

• Provide support for data warehouses, lakehouse architectures, feature stores, embedding indices, retrieval pipelines, and the infrastructure for model training, evaluation, and serving.

• Establish and advance observability practices across distributed systems, which encompass metrics, logging, tracing, and alerting.

• Implement reliability practices including SLOs/SLIs, incident response, postmortems, and on-call rotations.

• Design security and compliance measures for cloud infrastructure, secrets management, access control, and integration with industrial or legacy environments.

• Balance trade-offs between cost, latency, reliability, and developer velocity.

• Collaborate with engineering leadership to shape the infrastructure roadmap and platform strategy.

• Mentor engineers and promote operational excellence throughout the organization.

• Manage complex, high-stakes infrastructure challenges from start to finish, designing systems for future scalability.


⛳️ Requirements

• A minimum of 6 years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or roles focused on infrastructure software engineering.

• Extensive hands-on experience with AWS, GCP, or Azure at a production scale.

• Proven experience with Kubernetes in production settings, including the scheduling of GPU workloads.

• Proficiency in infrastructure-as-code tools such as Terraform, Pulumi, or similar alternatives.

• Familiarity with CI/CD systems like GitHub Actions, GitLab CI, CircleCI, Jenkins, or ArgoCD.

• Experience in designing and managing observability stacks such as Prometheus, Grafana, Datadog, or OpenTelemetry.

• Experience in supporting ML/AI infrastructure is highly desirable.

• Strong scripting/programming skills in Python, Go, or Bash.

• Proven capability to independently define and lead infrastructure projects from conception to production deployment.

• Strong instincts in incident management with the ability to lead outage responses and conduct root-cause analyses.

• Preferred: experience in integrating cloud and edge/on-premises environments, particularly within industrial or manufacturing contexts.

• Preferred: familiarity with Snowflake, BigQuery, Redshift, or Databricks.

• Preferred: knowledge of service mesh, zero-trust networking, or industrial compliance frameworks such as SOC 2 or IEC 62443.

• Preferred: experience in developing internal developer platforms or self-service infrastructure tools.

• Preferred: experience in scaling infrastructure teams or setting technical direction at the Staff level.


🏝️ Benefits

• Comprehensive salary and equity package.

• Significant opportunities for career development and advancement.

• Innovative environment centered on transforming heavy industries through AI and automation.

• Collaborative culture that values innovation, discipline, and continuous improvement.

People also viewed

Chalice AI14 hours ago

Senior Site Reliability Engineer

US flagNew York OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$180k – $190k/year
ApplyView job
DaVita Kidney Care14 hours ago

Senior DevOps Engineer

US flagTennessee OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$115k – $183k/year
ApplyView job
WEX16 hours ago

SRE & Application Services Intern – Graduate/Master's

US flagUnited States OnlyInternshipDevOps & Site Reliability Engineer (SRE)$30 – $45/hour
ApplyView job
WEX16 hours ago

DevOps Engineer, AI Engineering Intern

US flagUnited States OnlyInternshipDevOps & Site Reliability Engineer (SRE)
ApplyView job
Dscout16 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DaVita Kidney Care17 hours ago

Senior Manager, DevOps

US flagFlorida OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$115k – $183k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers