
Staff Site Reliability Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Europe.
• Design, deploy, and manage scalable and secure production environments (preferably AWS).
• Drive reliability enhancements across various engineering streams.
• Create and advance Kubernetes-based infrastructure, including migration and optimization efforts.
• Establish and uphold robust Infrastructure-as-Code standards.
• Identify and implement SLIs, SLOs, and error budgets.
• Enhance observability across applications, infrastructure, data pipelines, and machine learning systems.
• Collaborate closely with product and data teams to incorporate model analytics and product telemetry into reliability insights.
• Optimize and manage the entire CI/CD pipeline, from build to deployment to rollback.
• Increase release safety, deployment frequency, and predictability of SLAs.
• Lead incident response for intricate cross-system failures and facilitate postmortem analyses.
• Minimize operational toil through automation and advancements in platform engineering.
• Develop processes and tools to standardize, absorb, and troubleshoot customer environments.
• Support and operationalize machine learning workloads (implementing MLOps practices, including model deployment, monitoring, and retraining workflows).
• Ensure that infrastructure meets enterprise-grade security and regulatory standards.
• Mentor engineers and elevate the overall reliability standard across teams.
• Extensive hands-on experience in Site Reliability Engineering (SRE) or Production Engineering roles.
• Proven experience in building or scaling SRE practices in high-growth or complex settings.
• In-depth knowledge of AWS or Azure-based cloud infrastructure.
• Strong background in Kubernetes (including migration, scaling, and production hardening).
• Advanced experience with Infrastructure-as-Code (Terraform or similar tools).
• Comprehensive experience in designing and optimizing end-to-end CI/CD pipelines.
• Solid experience with observability tools across distributed systems.
• Proven ability to troubleshoot complex multi-tenant or customer-hosted environments.
• Experience in supporting production data platforms and machine learning systems.
• MLOps experience, including model deployment and monitoring.
• Strong understanding of distributed systems, scalability, and fault tolerance.
• A systems thinker who comprehends interactions across infrastructure, product, data, and machine learning.
• Exceptional communication skills with the ability to work collaboratively across functions.
• Health insurance
• Paid time off
• Professional development
CWILL
a37
GT
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.