Remotery

Senior Site Reliability Engineer

Posted Aug 5

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the reliability, scalability, and observability of essential financial SaaS applications and infrastructure.

• Design, implement, and sustain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical systems.

• Lead the development of monitoring, logging, and distributed tracing architecture while deploying observability tools.

• Create runbooks, incident response protocols, and post-incident review procedures.

• Provide mentorship to the team on incident management and conduct blameless postmortems.

• Architect and implement cloud infrastructure on AWS or Azure.

• Apply infrastructure-as-code practices, ensuring high availability, disaster recovery, and business continuity.

• Enhance automation and AIOps capabilities for incident detection, self-healing mechanisms, and intelligent alerting.

• Propel reliability improvements through load testing, chaos engineering, and failure scenario analyses.

• Collaborate with application and backend teams to ensure reliable system design, conduct architecture reviews, and perform reliability assessments.

• Develop production-grade Python tools for automation, metrics collection, alert management, and operational workflows.

• Advocate for infrastructure security and compliance within a regulated fintech environment.


⛳️ Requirements

• Over 7 years of experience in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with substantial responsibilities for production systems.

• Expert-level proficiency with Azure or AWS, including in-depth knowledge of compute, networking, storage, and managed services.

• Experience managing large-scale infrastructure.

• Proven expertise in designing and implementing monitoring, alerting, logging, and distributed tracing solutions.

• Practical experience with observability platforms such as Prometheus, Grafana, ELK, Datadog, or New Relic.

• Strong foundation in SLOs, SLIs, and SLAs; familiarity with error budgets.

• Experience in designing and troubleshooting highly available, resilient, and scalable systems.

• Comprehensive understanding of distributed systems concepts and potential failure mechanisms.

• Proficient in Python, PowerShell, bash, and similar scripting languages.

• Hands-on experience with AIOps practices, including event correlation, intelligent alerting, predictive analytics, and automated remediation.

• Familiarity with infrastructure-as-code tools such as Terraform, CloudFormation, or Ansible.

• Experience with version control and CI/CD pipeline design.

• Proven track record in incident management and on-call responsibilities.

• Exceptional communication skills with the ability to work cross-functionally and mentor junior engineers.

• Preferred: experience in fintech, payments, banking, or regulated industries.

• Preferred: experience with Kubernetes and container orchestration.

• Preferred: knowledge in observability as code, chaos engineering, open-source contributions, security hardening, and database optimization.


🏝️ Benefits

• Comprehensive insurance coverage including medical, dental, vision, life, and disability.

• Flexible paid time off.

• Paid holidays.

• 401(k) plan with company matching.

• Option for remote work.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers