Remotery

Staff Site Reliability Engineer

Posted Jul 28

This is a fully remote position, open to applicants in Europe.

📋 Description

• Design, deploy, and manage scalable and secure production environments (preferably AWS).

• Drive reliability enhancements across various engineering streams.

• Create and advance Kubernetes-based infrastructure, including migration and optimization efforts.

• Establish and uphold robust Infrastructure-as-Code standards.

• Identify and implement SLIs, SLOs, and error budgets.

• Enhance observability across applications, infrastructure, data pipelines, and machine learning systems.

• Collaborate closely with product and data teams to incorporate model analytics and product telemetry into reliability insights.

• Optimize and manage the entire CI/CD pipeline, from build to deployment to rollback.

• Increase release safety, deployment frequency, and predictability of SLAs.

• Lead incident response for intricate cross-system failures and facilitate postmortem analyses.

• Minimize operational toil through automation and advancements in platform engineering.

• Develop processes and tools to standardize, absorb, and troubleshoot customer environments.

• Support and operationalize machine learning workloads (implementing MLOps practices, including model deployment, monitoring, and retraining workflows).

• Ensure that infrastructure meets enterprise-grade security and regulatory standards.

• Mentor engineers and elevate the overall reliability standard across teams.


⛳️ Requirements

• Extensive hands-on experience in Site Reliability Engineering (SRE) or Production Engineering roles.

• Proven experience in building or scaling SRE practices in high-growth or complex settings.

• In-depth knowledge of AWS or Azure-based cloud infrastructure.

• Strong background in Kubernetes (including migration, scaling, and production hardening).

• Advanced experience with Infrastructure-as-Code (Terraform or similar tools).

• Comprehensive experience in designing and optimizing end-to-end CI/CD pipelines.

• Solid experience with observability tools across distributed systems.

• Proven ability to troubleshoot complex multi-tenant or customer-hosted environments.

• Experience in supporting production data platforms and machine learning systems.

• MLOps experience, including model deployment and monitoring.

• Strong understanding of distributed systems, scalability, and fault tolerance.

• A systems thinker who comprehends interactions across infrastructure, product, data, and machine learning.

• Exceptional communication skills with the ability to work collaboratively across functions.


🏝️ Benefits

• Health insurance

• Paid time off

• Professional development

People also viewed

CWILL19 hours ago

DevOps/SRE Engineer, Bilingual Mandarin

US flagCalifornia, +4 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$100k – $130k/year
ApplyView job
a3720 hours ago

Forward Deployed DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GT20 hours ago

Site Reliability Engineer, SRE

PL flagPoland, +2 more statesFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
Sigma Software Group20 hours ago

DevOps Engineer

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Applaudo21 hours ago

Google Cloud DevOps Engineer – Temporary Contract

CO flagColombia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Branch21 hours ago

Cloud Operations Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers