
Software Engineer III – Site Reliability
Posted Jul 19

Posted Jul 19
This is a fully remote position, open to applicants in United States.
• Take ownership and enhance our SLI/SLO and error-budget frameworks.
• Lead incident response efforts and facilitate postmortem analyses.
• Develop and sustain observability across metrics, logs, and traces using Datadog.
• Design and manage resilient, scalable infrastructure leveraging Infrastructure as Code with Terraform.
• Oversee production Kubernetes and container workloads.
• Manage CI/CD pipelines along with secure deployment strategies.
• Maintain the security controls embedded within the delivery pipeline.
• Implement and uphold policy-as-code practices.
• Lead vulnerability triage and remediation service level agreements (SLAs) for findings at both the pipeline and infrastructure levels.
• Collaborate with our Security Engineer and the wider Security & Reliability teams.
• Engage in and enhance the on-call rotation.
• Mentor team members and engineers across the organization on reliability patterns and operational best practices.
• Minimum of 5 years in site reliability, platform, or infrastructure engineering, demonstrating clear senior-level ownership of production systems.
• Proficient programming skills for automation and tooling in languages such as Go, Python, Typescript, or similar.
• Extensive hands-on experience with a major cloud platform (AWS preferred), Kubernetes, and Infrastructure as Code (Terraform preferred).
• Proven experience in leading incident responses and establishing SLO-driven reliability practices.
• Proficient use of observability tools (Datadog preferred).
• Practical experience in integrating security into CI/CD pipelines, including SAST/DAST/SCA tooling, dependency scanning, or policy-as-code.
• Strong comprehension of cloud security fundamentals.
• Ability to exercise judgment and communicate effectively to raise security or reliability concerns with senior engineers.
• Familiarity with policy-as-code frameworks (especially Kyverno, though tools like OPA/Rego or Conftest are also relevant).
• Experience in regulated or compliance-oriented environments (such as SOC 2, PCI DSS, HIPAA) is a plus.
• Knowledge of chaos engineering or participation in game-day scenarios is advantageous.
• Experience in supporting B2C/mobile backend environments characterized by high traffic, rapid iteration, and stringent reliability requirements is a plus.
• Healthcare coverage.
• Parental planning support.
• Mental health benefits.
• Annual performance bonus.
• 401(k) plan with matching contributions.
• Responsible time off policy.
• Monthly wellness and technology allowances.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.