
Staff DevOps Engineer
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in California.
• Take ownership of and enhance the fundamental infrastructure encompassing compute, networking, storage, and deployment systems.
• Design and manage CI/CD pipelines for AI, data, and product engineering teams.
• Develop and sustain infrastructure-as-code for consistent and auditable cloud, on-premises, and edge environments.
• Architect and oversee Kubernetes platforms tailored for training, inference, and application workloads, including GPU scheduling and autoscaling capabilities.
• Provide support for data warehouses, lakehouse architectures, feature stores, embedding indices, retrieval pipelines, and the infrastructure for model training, evaluation, and serving.
• Establish and advance observability practices across distributed systems, which encompass metrics, logging, tracing, and alerting.
• Implement reliability practices including SLOs/SLIs, incident response, postmortems, and on-call rotations.
• Design security and compliance measures for cloud infrastructure, secrets management, access control, and integration with industrial or legacy environments.
• Balance trade-offs between cost, latency, reliability, and developer velocity.
• Collaborate with engineering leadership to shape the infrastructure roadmap and platform strategy.
• Mentor engineers and promote operational excellence throughout the organization.
• Manage complex, high-stakes infrastructure challenges from start to finish, designing systems for future scalability.
• A minimum of 6 years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or roles focused on infrastructure software engineering.
• Extensive hands-on experience with AWS, GCP, or Azure at a production scale.
• Proven experience with Kubernetes in production settings, including the scheduling of GPU workloads.
• Proficiency in infrastructure-as-code tools such as Terraform, Pulumi, or similar alternatives.
• Familiarity with CI/CD systems like GitHub Actions, GitLab CI, CircleCI, Jenkins, or ArgoCD.
• Experience in designing and managing observability stacks such as Prometheus, Grafana, Datadog, or OpenTelemetry.
• Experience in supporting ML/AI infrastructure is highly desirable.
• Strong scripting/programming skills in Python, Go, or Bash.
• Proven capability to independently define and lead infrastructure projects from conception to production deployment.
• Strong instincts in incident management with the ability to lead outage responses and conduct root-cause analyses.
• Preferred: experience in integrating cloud and edge/on-premises environments, particularly within industrial or manufacturing contexts.
• Preferred: familiarity with Snowflake, BigQuery, Redshift, or Databricks.
• Preferred: knowledge of service mesh, zero-trust networking, or industrial compliance frameworks such as SOC 2 or IEC 62443.
• Preferred: experience in developing internal developer platforms or self-service infrastructure tools.
• Preferred: experience in scaling infrastructure teams or setting technical direction at the Staff level.
• Comprehensive salary and equity package.
• Significant opportunities for career development and advancement.
• Innovative environment centered on transforming heavy industries through AI and automation.
• Collaborative culture that values innovation, discipline, and continuous improvement.
Chalice AI
DaVita Kidney Care
WEX
WEX
Get handpicked remote jobs straight to your inbox weekly.