
Site Reliability Engineer
Posted Jul 26

Posted Jul 26
This is a fully remote position, open to applicants in United States.
• Take ownership of and enhance our SLI/SLO and error-budget frameworks, utilizing them to guide prioritization and product decisions.
• Lead the incident response process, facilitate postmortems, and convert findings into systemic improvements instead of temporary solutions.
• Develop and sustain observability across metrics, logs, and traces (Datadog), enhancing signal and minimizing alert fatigue.
• Design and manage resilient, scalable infrastructure using Infrastructure as Code (Terraform).
• Oversee production Kubernetes and container workloads, including capacity planning and optimizing cloud costs.
• Manage CI/CD pipelines and implement safe deployment strategies (canary, progressive rollout, fast rollback).
• Be responsible for the security controls within the delivery pipeline — integrating and fine-tuning SAST, DAST, and SCA scanning (for instance, in GitHub Actions) to ensure issues are identified while the code is under review.
• Implement and uphold policy-as-code (for example, OPA/Rego, Kyverno, or Conftest) to prevent unsafe infrastructure and Kubernetes changes at admission time.
• Lead vulnerability triage and remediation SLAs for pipeline and infrastructure-level findings, prioritizing based on actual risk.
• Collaborate with our Security Engineer and the wider Security & Reliability teams — you are responsible for security in the pipeline while working together on the broader aspects, avoiding duplication of efforts.
• Engage in and enhance the on-call rotation; create runbooks and automation tools that make on-call responsibilities sustainable.
• Mentor team members and engineers throughout the organization on reliability patterns and operational best practices.
• Over 5 years of experience in site reliability, platform, or infrastructure engineering.
• Proficient programming skills for automation and tooling (Go, Python, Typescript, or similar).
• Extensive hands-on experience with a major cloud platform (AWS is advantageous), Kubernetes, and Infrastructure as Code (Terraform is beneficial).
• A proven history of leading incident response and establishing SLO-driven reliability practices.
• Working knowledge of observability tools (Datadog is a plus).
• Practical experience in integrating security within CI/CD pipelines — SAST/DAST/SCA tooling, dependency scanning, or policy-as-code.
• Strong understanding of cloud security fundamentals (identity/IAM, least-privilege patterns, policy/guardrails, secrets management).
• The ability to communicate effectively and exercise sound judgment when escalating a security or reliability issue with a senior engineer, framing it as a collaborative problem to solve rather than a conflict.
• Familiarity with policy-as-code frameworks (especially Kyverno, but tools like OPA/Rego or Conftest are also relevant) enforced at admission time is a plus.
• Experience in regulated or compliance-driven environments (SOC 2, PCI DSS, HIPAA) is a plus.
• Experience with chaos engineering or game-day scenarios is beneficial.
• Background in supporting B2C/mobile backend environments with high traffic, rapid iteration, and strong reliability requirements is a plus.
• Healthcare coverage.
• Parental planning support.
• Mental health benefits.
• Annual performance bonuses.
• A 401(k) plan with matching contributions.
• Responsible time off policy.
• Monthly allowances for wellness and technology.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.