Staff Site Reliability Engineer

Posted 4 days ago

This is a fully remote position, open to applicants in Spain.

📋 Description

• Take ownership of enhancing the reliability, scalability, and operational efficiency of enterprise data and AI platforms.

• Establish, measure, and enforce Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets across platform services.

• Minimize Mean Time to Detect (MTTD) and Mean Time to Recovery (MTTR) through improved visibility, automated alerts, and runbook-driven incident responses.

• Lead constructive post-mortem analyses and implement sustainable reliability enhancements.

• Identify, assess, and mitigate operational inefficiencies; monitor toil percentage and impose capacity limits.

• Develop and sustain self-service infrastructure capabilities for product and data engineering teams.

• Automate deployment pipelines, configuration management, and operational workflows by utilizing Infrastructure-as-Code, Terraform, Helm, and GitOps methodologies.

• Design and optimize workloads utilizing GCP services such as GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI.

• Oversee architecture decisions spanning multi-cloud environments, including GCP and Azure.

• Craft self-healing infrastructure, auto-scaling strategies, and capacity planning models.

• Advocate for data security and governance practices, encompassing encryption, least-privilege IAM, secrets management, and audit logging.

• Implement Site Reliability Engineering (SRE) principles to agentic AI workloads and establish reliability expectations for LLM-based and multi-agent systems.

• Collaborate with AI Platform teams to operationalize agentic pipelines, incorporating monitoring, drift detection, and rollback functionalities.

• Promote the utilization of AI-assisted operational tools for observability, anomaly detection, and predictive incident management.

• Mentor junior and mid-level SREs through code evaluations and reliability assessments.

• Work in conjunction with Data Engineering, Software Engineering, Security, and Product teams.

• Effectively communicate platform health, risk posture, and reliability strategies to both technical and executive audiences.


⛳️ Requirements

• Master's degree in Computer Science, Engineering, or a related discipline (or equivalent practical experience).

• Over 10 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering, with a proven history in large-scale production settings.

• Extensive hands-on knowledge of GCP, particularly GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI.

• Proficient in Infrastructure-as-Code tools (Terraform, Helm) and GitOps workflows (ArgoCD, Flux).

• Strong expertise in observability tools, including distributed tracing, structured logging, metrics pipelines, and alerting systems such as Cloud Monitoring, Datadog, and Prometheus/Grafana.

• Demonstrated experience in defining and operating against SLOs, SLIs, and error budgets within production environments.

• Comprehensive understanding of data security principles: IAM, encryption, secrets management, network policies, and compliance frameworks.

• Experience with multi-cloud environments (GCP + Azure), focusing on cross-cloud networking, identity federation, and cost management.

• Proven capability to lead without formal authority, influence cross-team engineers, and drive organization-wide reliability enhancements.

• Exceptional written and verbal communication skills; able to convey complex reliability concepts to non-technical stakeholders.

• Proficient in at least one systems or scripting language: Python, Go, or Bash.

• Experience with container orchestration technologies (Kubernetes/GKE), service mesh, and traffic management patterns.

• Familiarity with data engineering practices: batch and streaming pipelines, data warehouses, and the operational challenges associated with large-scale data platforms.

• Understanding of agentic AI architectures and the reliability challenges faced by LLM-based, event-driven, and multi-agent systems.

• Working knowledge of data governance standards, data lineage tools, and platform-level data quality enforcement.

• Experience managing data platforms within retail or e-commerce settings.

• Familiarity with SRE principles in action, including error budget policies, CRE engagements, and production readiness assessments.

• Exposure to FinOps practices, such as cloud cost attribution, optimization of commitments, and unit economics related to data workloads.

• Experience with Azure-native services (AKS, Azure Data Factory, Event Hubs) and cross-cloud identity management.

• Previous experience in a Staff or Principal-level SRE role with an organization-wide scope.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health benefits including medical, dental, and vision coverage.

• Flexible work arrangements and remote work options.

• Professional development opportunities and support for continuous learning.

• Collaborative and inclusive work environment.

People also viewed

Horizon3.ai21 hours ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH21 hours ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM21 hours ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies21 hours ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)21 hours ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC21 hours ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers