
Staff Site Reliability Engineer
Posted Sep 14

Posted Sep 14
This is a fully remote position, open to applicants in United States.
• Act as the first dedicated Site Reliability Engineer at Fingerprint
• Collaborate with the Architect on platform design and with the Cloud Platform team on infrastructure
• Work alongside every product team to manage their systems effectively
• Establish SLIs and SLOs for essential request paths, ensuring they are visible and actionable
• Introduce and mentor teams on the concept of error budgets
• Oversee the reliability metrics utilized by leadership
• Enhance incident detection, response, communication, postmortems, and follow-up processes
• Improve the quality of alerts, anomaly detection, escalation design, and shared tools
• Conduct reliability reviews for high-risk changes and new services
• Initiate game days and chaos engineering exercises
• Engage with teams through time-limited reliability projects
• Cultivate Staff and Lead engineers to become leaders in reliability
• Formalize practices related to production-readiness, on-call procedures, runbooks, and change safety
• Collaborate with the Architect and technical leads to integrate reliability into systems
• Investigate production incidents and create tools, dashboards, and reference implementations
• Spearhead AI integration for incident investigation, postmortems, runbooks, observability, and safe AI-assisted operations
• Report directly to the VP of Engineering
• Over 10 years of engineering experience
• At least 3 years in an SRE, production engineer, or reliability-focused Staff engineer role across multiple teams
• Proven experience in owning reliability for a platform
• Extensive knowledge of SLI/SLO design and practical implementation of error budgets
• Experience in promoting reliability best practices within product teams
• Proven track record in leading incident response and postmortems for high-severity, customer-facing incidents
• Hands-on expertise with distributed-systems failure modes, including cache/database saturation, cascading failures, retry storms, capacity limits, degradation, and load shedding
• Experience in high-throughput, low-latency environments
• Proficient in Kubernetes, AWS, and modern observability tools like Datadog or similar
• Ability to read and write production code in Go, TypeScript, or comparable languages
• Familiarity with infrastructure as code
• Proven ability to lead through influence across teams
• Experience coaching engineers to take ownership of reliability
• Excellent written communication skills and documented decision-making processes
• Regular utilization of AI tools for incident investigation, telemetry analysis, runbooks, postmortems, and tooling
• Ability to organize operational data for secure use by both humans and AI agents
• A pragmatic approach to balancing reliability, delivery, and risk
• Must have authorization to work from the specified home location
• Nice to have: experience in fraud detection, identity, payments, or other adversarial real-time domains
• Nice to have: experience with multi-region, cell-based, or failure-isolation architecture
• Nice to have: familiarity with Elasticsearch, Redis, DynamoDB, or Kafka at scale
• Nice to have: knowledge of FinOps and cloud infrastructure reliability/cost trade-offs
• Pay transparency
• Fully remote work arrangement
• Flexibility to work from almost any country, subject to local restrictions
• Inclusive work environment
• No visa sponsorship required; teammates can work from their authorized home locations
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.