Staff Site Reliability Engineer

Posted Sep 14

This is a fully remote position, open to applicants in United States.

📋 Description

• Act as the first dedicated Site Reliability Engineer at Fingerprint

• Collaborate with the Architect on platform design and with the Cloud Platform team on infrastructure

• Work alongside every product team to manage their systems effectively

• Establish SLIs and SLOs for essential request paths, ensuring they are visible and actionable

• Introduce and mentor teams on the concept of error budgets

• Oversee the reliability metrics utilized by leadership

• Enhance incident detection, response, communication, postmortems, and follow-up processes

• Improve the quality of alerts, anomaly detection, escalation design, and shared tools

• Conduct reliability reviews for high-risk changes and new services

• Initiate game days and chaos engineering exercises

• Engage with teams through time-limited reliability projects

• Cultivate Staff and Lead engineers to become leaders in reliability

• Formalize practices related to production-readiness, on-call procedures, runbooks, and change safety

• Collaborate with the Architect and technical leads to integrate reliability into systems

• Investigate production incidents and create tools, dashboards, and reference implementations

• Spearhead AI integration for incident investigation, postmortems, runbooks, observability, and safe AI-assisted operations

• Report directly to the VP of Engineering


⛳️ Requirements

• Over 10 years of engineering experience

• At least 3 years in an SRE, production engineer, or reliability-focused Staff engineer role across multiple teams

• Proven experience in owning reliability for a platform

• Extensive knowledge of SLI/SLO design and practical implementation of error budgets

• Experience in promoting reliability best practices within product teams

• Proven track record in leading incident response and postmortems for high-severity, customer-facing incidents

• Hands-on expertise with distributed-systems failure modes, including cache/database saturation, cascading failures, retry storms, capacity limits, degradation, and load shedding

• Experience in high-throughput, low-latency environments

• Proficient in Kubernetes, AWS, and modern observability tools like Datadog or similar

• Ability to read and write production code in Go, TypeScript, or comparable languages

• Familiarity with infrastructure as code

• Proven ability to lead through influence across teams

• Experience coaching engineers to take ownership of reliability

• Excellent written communication skills and documented decision-making processes

• Regular utilization of AI tools for incident investigation, telemetry analysis, runbooks, postmortems, and tooling

• Ability to organize operational data for secure use by both humans and AI agents

• A pragmatic approach to balancing reliability, delivery, and risk

• Must have authorization to work from the specified home location

• Nice to have: experience in fraud detection, identity, payments, or other adversarial real-time domains

• Nice to have: experience with multi-region, cell-based, or failure-isolation architecture

• Nice to have: familiarity with Elasticsearch, Redis, DynamoDB, or Kafka at scale

• Nice to have: knowledge of FinOps and cloud infrastructure reliability/cost trade-offs


🏝️ Benefits

• Pay transparency

• Fully remote work arrangement

• Flexibility to work from almost any country, subject to local restrictions

• Inclusive work environment

• No visa sponsorship required; teammates can work from their authorized home locations

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF13 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam17 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers