
Staff Site Reliability Engineer
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Spain.
• Take ownership of enhancing the reliability, scalability, and operational efficiency of enterprise data and AI platforms.
• Establish, measure, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets across platform services.
• Minimize Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR) through improved visibility, automated alerts, and runbook-driven incident responses.
• Lead constructive post-mortem analyses and implement sustainable reliability enhancements.
• Identify, assess, and mitigate operational inefficiencies; monitor toil percentage and impose capacity limits.
• Develop and sustain self-service infrastructure capabilities for product and data engineering teams.
• Automate deployment pipelines, configuration management, and operational workflows by utilizing Infrastructure-as-Code, Terraform, Helm, and GitOps methodologies.
• Design and optimize workloads utilizing GCP services such as GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI.
• Oversee architecture decisions spanning multi-cloud environments, including GCP and Azure.
• Craft self-healing infrastructure, auto-scaling strategies, and capacity planning models.
• Advocate for data security and governance practices, encompassing encryption, least-privilege IAM, secrets management, and audit logging.
• Implement Site Reliability Engineering (SRE) principles to agentic AI workloads and establish reliability expectations for LLM-based and multi-agent systems.
• Collaborate with AI Platform teams to operationalize agentic pipelines, incorporating monitoring, drift detection, and rollback functionalities.
• Promote the utilization of AI-assisted operational tools for observability, anomaly detection, and predictive incident management.
• Mentor junior and mid-level SREs through code evaluations and reliability assessments.
• Work in conjunction with Data Engineering, Software Engineering, Security, and Product teams.
• Effectively communicate platform health, risk posture, and reliability strategies to both technical and executive audiences.
• Master's degree in Computer Science, Engineering, or a related discipline (or equivalent practical experience).
• Over 10 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a proven history in large-scale production settings.
• Extensive hands-on knowledge of GCP, particularly GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI.
• Proficient in Infrastructure-as-Code tools (Terraform, Helm) and GitOps workflows (ArgoCD, Flux).
• Strong expertise in observability tools, including distributed tracing, structured logging, metrics pipelines, and alerting systems such as Cloud Monitoring, Datadog, and Prometheus/Grafana.
• Demonstrated experience in defining and operating against SLOs, SLIs, and error budgets within production environments.
• Comprehensive understanding of data security principles: IAM, encryption, secrets management, network policies, and compliance frameworks.
• Experience with multi-cloud environments (GCP + Azure), focusing on cross-cloud networking, identity federation, and cost management.
• Proven capability to lead without formal authority, influence cross-team engineers, and drive organization-wide reliability enhancements.
• Exceptional written and verbal communication skills; able to convey complex reliability concepts to non-technical stakeholders.
• Proficient in at least one systems or scripting language: Python, Go, or Bash.
• Experience with container orchestration technologies (Kubernetes/GKE), service mesh, and traffic management patterns.
• Familiarity with data engineering practices: batch and streaming pipelines, data warehouses, and the operational challenges associated with large-scale data platforms.
• Understanding of agentic AI architectures and the reliability challenges faced by LLM-based, event-driven, and multi-agent systems.
• Working knowledge of data governance standards, data lineage tools, and platform-level data quality enforcement.
• Experience managing data platforms within retail or e-commerce settings.
• Familiarity with SRE principles in action, including error budget policies, CRE engagements, and production readiness assessments.
• Exposure to FinOps practices, such as cloud cost attribution, optimization of commitments, and unit economics related to data workloads.
• Experience with Azure-native services (AKS, Azure Data Factory, Event Hubs) and cross-cloud identity management.
• Previous experience in a Staff or Principal-level SRE role with an organization-wide scope.
• Competitive salary and performance-based bonuses.
• Comprehensive health benefits including medical, dental, and vision coverage.
• Flexible work arrangements and remote work options.
• Professional development opportunities and support for continuous learning.
• Collaborative and inclusive work environment.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.