SRE Specialist I

Posted 4 days ago

This is a fully remote position, open to applicants in California.

📋 Description

• Act as the technical SRE liaison for Identity & Fraud products, assisting both development and operations teams.

• Establish, implement, and oversee SLIs, SLOs, and SLAs that align with business goals.

• Lead analyses of incidents and suggest preventive and corrective measures to prevent future occurrences.

• Automate processes for provisioning, deployment, scaling, and recovery from failures.

• Design and uphold observability solutions encompassing logs, metrics, traces, and alerts.

• Engage in capacity and performance engineering activities.

• Assist in the architectural evolution with a focus on the resilience, scalability, and security of Identity & Fraud systems.

• Advocate for best practices in infrastructure as code, CI/CD, and safe change management.

• Support essential operations by diagnosing and resolving issues in real-time.


⛳️ Requirements

• Significant experience in SRE, DevOps, or Production Engineering within mission-critical environments.

• Proficiency in Kubernetes, Docker, and various cloud platforms (AWS, OCI, Azure, and GCP).

• Advanced expertise in automation and infrastructure as code tools (Terraform, Ansible, etc.).

• Familiarity with monitoring and observability tools, particularly Datadog, along with Prometheus, ELK, and Grafana.

• Experience with CI/CD pipelines and best practices related to version control and deployment.

• Strong skills in analyzing performance, troubleshooting, and optimizing distributed systems.

• Knowledge of both relational and non-relational databases.

• Capability to collaborate effectively with development, product, and operations teams.

• Excellent communication skills, systems thinking, and a focus on addressing complex challenges.

• Experience with resilience in identity and fraud systems.

• Relevant cloud certifications (AWS, OCI, Azure, or GCP).

• Experience in chaos engineering and resilience testing.

• Understanding of application and infrastructure security principles.


🏝️ Benefits

• Employees have the option to work remotely.

• A diverse, purpose-driven working environment.

• Opportunities for career advancement and professional development.

• A supportive environment that encourages a balance between career and personal commitments.

• Initiatives focused on employee well-being.

• Recognition and external certifications from Great Place To Work™ and Top Employers.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers