
Senior Site Reliability Engineer
Posted Aug 5

Posted Aug 5
This is a fully remote position, open to applicants in United States.
• Take ownership of the reliability, scalability, and observability of essential financial SaaS applications and infrastructure.
• Design, implement, and sustain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical systems.
• Lead the development of monitoring, logging, and distributed tracing architecture while deploying observability tools.
• Create runbooks, incident response protocols, and post-incident review procedures.
• Provide mentorship to the team on incident management and conduct blameless postmortems.
• Architect and implement cloud infrastructure on AWS or Azure.
• Apply infrastructure-as-code practices, ensuring high availability, disaster recovery, and business continuity.
• Enhance automation and AIOps capabilities for incident detection, self-healing mechanisms, and intelligent alerting.
• Propel reliability improvements through load testing, chaos engineering, and failure scenario analyses.
• Collaborate with application and backend teams to ensure reliable system design, conduct architecture reviews, and perform reliability assessments.
• Develop production-grade Python tools for automation, metrics collection, alert management, and operational workflows.
• Advocate for infrastructure security and compliance within a regulated fintech environment.
• Over 7 years of experience in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with substantial responsibilities for production systems.
• Expert-level proficiency with Azure or AWS, including in-depth knowledge of compute, networking, storage, and managed services.
• Experience managing large-scale infrastructure.
• Proven expertise in designing and implementing monitoring, alerting, logging, and distributed tracing solutions.
• Practical experience with observability platforms such as Prometheus, Grafana, ELK, Datadog, or New Relic.
• Strong foundation in SLOs, SLIs, and SLAs; familiarity with error budgets.
• Experience in designing and troubleshooting highly available, resilient, and scalable systems.
• Comprehensive understanding of distributed systems concepts and potential failure mechanisms.
• Proficient in Python, PowerShell, bash, and similar scripting languages.
• Hands-on experience with AIOps practices, including event correlation, intelligent alerting, predictive analytics, and automated remediation.
• Familiarity with infrastructure-as-code tools such as Terraform, CloudFormation, or Ansible.
• Experience with version control and CI/CD pipeline design.
• Proven track record in incident management and on-call responsibilities.
• Exceptional communication skills with the ability to work cross-functionally and mentor junior engineers.
• Preferred: experience in fintech, payments, banking, or regulated industries.
• Preferred: experience with Kubernetes and container orchestration.
• Preferred: knowledge in observability as code, chaos engineering, open-source contributions, security hardening, and database optimization.
• Comprehensive insurance coverage including medical, dental, vision, life, and disability.
• Flexible paid time off.
• Paid holidays.
• 401(k) plan with company matching.
• Option for remote work.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.