
Senior Site Reliability Engineer – Volcano
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Define and oversee service level objectives (SLOs), error budgets, and incident response strategies for all Volcano services.
• Create and implement multi-region Kubernetes infrastructure, including networking and data plane components.
• Set up deployment automation, canary release pipelines, and preview environment provisioning utilizing ArgoCD, Helm, and Terraform/Terragrunt.
• Design, manage, and secure multi-tenant PostgreSQL clusters, Redis caching layers, and object storage solutions.
• Instrument Volcano services with relevant service level indicators (SLIs) and develop dashboards, alerts, and runbooks using Datadog, Prometheus, and Grafana.
• Work collaboratively with the OCTO team, product engineering, and security to incorporate reliability and compliance into the system architecture.
• Assess architectural alternatives for edge runtimes, serverless computing, vector databases, and AI-native infrastructure components.
• Bachelor’s degree in Computer Science or a related field.
• Significant experience at the Staff or Principal individual contributor level in Site Reliability Engineering (SRE) or Platform Engineering.
• Demonstrated success in establishing SRE or platform engineering practices for developer-facing platforms or PaaS/SaaS solutions.
• Expertise in Kubernetes, including multi-tenant cluster design, networking (CNI, service mesh, ingress), autoscaling, and security enhancements.
• Healthcare benefits.
• 401(k) plan.
• Short- and long-term disability benefits.
• Basic life and AD&D insurance.
• Additional rewards for eligible roles, including sales incentives where applicable.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.