
Site Reliability Engineer
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in United States.
• Become the first dedicated Site Reliability Engineer on the Engineering team.
• Define the criteria for reliability in production systems and initiate the SRE practice from the ground up.
• Establish SLIs and SLOs for essential services.
• Create Grafana dashboards and implement alerting based on burn rates.
• Build, manage, and continuously enhance incident management processes, including PagerDuty, on-call training, and incident command.
• Oversee the observability stack comprehensively, covering metrics, logs, traces, RUM, and synthetic checks.
• Collaborate with engineering teams to refine SLIs, SLOs, and error budgets while coaching them on best practices in SRE and observability.
• Automate repetitive operational tasks through infrastructure as code and tooling.
• Design and execute load/performance tests and chaos engineering game days.
• Independently lead reliability and infrastructure projects, moving at a startup pace.
• Create documented incident management processes from detection to blameless postmortems.
• Minimize alert noise and MTTR, while enhancing confidence in reliability signals for release and investment decisions.
• A Bachelor's degree in Computer Science, Computer Engineering, or a similar field, with a strong foundation in operating systems, databases, and networking; a rigorous equivalent is acceptable.
• A solid understanding of systems is essential for diagnosing new failures.
• Daily use of AI tools for writing and debugging code, creating dashboards and alerts, and enhancing execution speed.
• Proven experience in independently managing large, ambiguous reliability or infrastructure projects to completion, generally requiring 5+ years in an SRE, DevOps, or production engineering position.
• Practical application of SLI/SLO/error-budget methodologies.
• Significant experience with Grafana and PromQL.
• Familiarity with Grafana Alloy for Loki logs and metrics backends like Prometheus or Datadog.
• Experience with OpenTelemetry and tracing/APM backends such as SigNoz, Uptrace, Tempo, Datadog, or New Relic.
• Experience with Real User Monitoring (RUM) and synthetic monitoring tools, such as Grafana Faro, Grafana Synthetic Monitoring, or k6.
• Capability to design on-call rotations and incident command practices utilizing PagerDuty or similar tools.
• Hands-on experience with load/performance testing frameworks like Locust, k6, or JMeter.
• Familiarity with chaos engineering practices.
• Proficient in Python, Go, or Bash.
• Practical experience with Infrastructure as Code tools such as Terraform or Ansible.
• Hands-on experience with Kubernetes.
• Knowledge of at least one major cloud platform: AWS, GCP, or Azure.
• Ability to effectively communicate technical root causes, trade-offs, implementation details, mitigations, fixes, reliability status, risks, and priorities to both technical and business stakeholders.
• Capability to work independently, swiftly resolve ambiguous issues, take ownership, and manage multiple tasks under time constraints.
• Comprehensive growth opportunities for personal and professional development.
• Dynamic and collaborative work environment.
• Excellent work/life balance.
• Medical insurance.
• Dental insurance.
• Vision insurance.
• Life insurance.
• 401k matching.
• Paid time off (PTO).
• Two office locations: Downtown Atlanta and Halcyon in Alpharetta.
FourEnergy GmbH
ICF
Mastercam
C&S Informática
Get handpicked remote jobs straight to your inbox weekly.