Site Reliability Engineer, Observabilidad IA

Posted 4 days ago

This is a fully remote position, open to applicants in Spain.

📋 Description

• Ensure the reliability, availability, and performance of AI platforms and products in production environments.

• Define and manage SLIs, SLOs, and Error Budgets, collaborating with product teams to balance innovation and stability.

• Design and drive observability and monitoring solutions, particularly in Artificial Intelligence environments.

• Lead incident management and coordinate problem resolution in critical settings.

• Foster a culture of continuous improvement.

• Automate operational processes and develop self-remediation mechanisms to enhance efficiency and reduce manual tasks.

• Collaborate with development, platform, and AI teams to ensure the scalability, resilience, and optimization of services.


⛳️ Requirements

• 3-6 years of experience as a Site Reliability Engineer (SRE), DevOps Engineer, Production Engineer, or Observability Engineer in production environments.

• Experience working with AWS and/or Azure.

• Proficiency with Kubernetes-based platforms.

• Familiarity with observability tools such as Prometheus, Grafana, and OpenTelemetry.

• Knowledge of Infrastructure as Code using Terraform and/or Bicep.

• Experience in CI/CD environments.

• Proficient in Python with a focus on automation and process improvement.

• Experience in Machine Learning, LLMs, GPUs, or AI platforms is a plus.

• English language proficiency at B2/C1 level.


🏝️ Benefits

• Immediate incorporation into a leading company in the IT sector with a high degree of expertise in Data & Analytics and ongoing expansion.

• Job stability through a permanent contract.

• Extensive opportunities for professional development and growth within the company.

• 100% remote work model from anywhere in Spain.

• Highly competitive compensation package in line with the candidate's value.

• Flexible compensation plans: restaurant card, transportation card, and childcare card.

• Health insurance.

• GYMPASS.

• Tailored training plans for each profile: technical courses, official certifications, and language training.

• Special discount portal for employees.

• Positive work environment and highly collaborative atmosphere.

People also viewed

FourEnergy GmbH13 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF15 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam19 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática1 day ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers