Site Reliability Engineer

Posted Aug 31

This is a fully remote position, open to applicants in United States.

📋 Description

• Become the first dedicated Site Reliability Engineer on the Engineering team.

• Define the criteria for reliability in production systems and initiate the SRE practice from the ground up.

• Establish SLIs and SLOs for essential services.

• Create Grafana dashboards and implement alerting based on burn rates.

• Build, manage, and continuously enhance incident management processes, including PagerDuty, on-call training, and incident command.

• Oversee the observability stack comprehensively, covering metrics, logs, traces, RUM, and synthetic checks.

• Collaborate with engineering teams to refine SLIs, SLOs, and error budgets while coaching them on best practices in SRE and observability.

• Automate repetitive operational tasks through infrastructure as code and tooling.

• Design and execute load/performance tests and chaos engineering game days.

• Independently lead reliability and infrastructure projects, moving at a startup pace.

• Create documented incident management processes from detection to blameless postmortems.

• Minimize alert noise and MTTR, while enhancing confidence in reliability signals for release and investment decisions.


⛳️ Requirements

• A Bachelor's degree in Computer Science, Computer Engineering, or a similar field, with a strong foundation in operating systems, databases, and networking; a rigorous equivalent is acceptable.

• A solid understanding of systems is essential for diagnosing new failures.

• Daily use of AI tools for writing and debugging code, creating dashboards and alerts, and enhancing execution speed.

• Proven experience in independently managing large, ambiguous reliability or infrastructure projects to completion, generally requiring 5+ years in an SRE, DevOps, or production engineering position.

• Practical application of SLI/SLO/error-budget methodologies.

• Significant experience with Grafana and PromQL.

• Familiarity with Grafana Alloy for Loki logs and metrics backends like Prometheus or Datadog.

• Experience with OpenTelemetry and tracing/APM backends such as SigNoz, Uptrace, Tempo, Datadog, or New Relic.

• Experience with Real User Monitoring (RUM) and synthetic monitoring tools, such as Grafana Faro, Grafana Synthetic Monitoring, or k6.

• Capability to design on-call rotations and incident command practices utilizing PagerDuty or similar tools.

• Hands-on experience with load/performance testing frameworks like Locust, k6, or JMeter.

• Familiarity with chaos engineering practices.

• Proficient in Python, Go, or Bash.

• Practical experience with Infrastructure as Code tools such as Terraform or Ansible.

• Hands-on experience with Kubernetes.

• Knowledge of at least one major cloud platform: AWS, GCP, or Azure.

• Ability to effectively communicate technical root causes, trade-offs, implementation details, mitigations, fixes, reliability status, risks, and priorities to both technical and business stakeholders.

• Capability to work independently, swiftly resolve ambiguous issues, take ownership, and manage multiple tasks under time constraints.


🏝️ Benefits

• Comprehensive growth opportunities for personal and professional development.

• Dynamic and collaborative work environment.

• Excellent work/life balance.

• Medical insurance.

• Dental insurance.

• Vision insurance.

• Life insurance.

• 401k matching.

• Paid time off (PTO).

• Two office locations: Downtown Atlanta and Halcyon in Alpharetta.

People also viewed

FourEnergy GmbH11 hours ago

Senior DevOps Engineer – Operations

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
ICF14 hours ago

Lead DevOps Engineer

US flagVirginia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$131.3k – $223.1k/year
ApplyView job
Mastercam18 hours ago

DevSecOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
C&S Informática22 hours ago

DevOps Engineer – Freelance/Contract, Mid-Level/Senior

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Convene1 day ago

Support and Deployment Engineer

SA flagSaudi Arabia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Verity Group1 day ago

SRE Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers