
Site Reliability Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Texas.
• Define and implement Service Level Objectives (SLOs) that are in alignment with customer SLAs.
• Develop and sustain the observability stack encompassing metrics, logs, traces, and alerting systems.
• Lead the incident response efforts and facilitate post-incident review meetings.
• Propel automation initiatives to minimize toil and enhance mean-time-to-recover (MTTR).
• Create and maintain operational runbooks in collaboration with the NOC.
• Oversee the on-call rotation, escalation paths, and tools for incident management.
• Collaborate across departments with NOC, Platform Engineering, and Network Engineering teams.
• Lead chaos engineering efforts, organize game days, and manage reliability testing programs.
• Generate SLA performance reports in partnership with the SLA Manager.
• Mentor junior engineers and foster a positive engineering culture.
• A minimum of 5 years of experience in SRE, DevOps, or production engineering roles.
• Proficient programming skills in Go, Python, or both.
• Practical experience in managing Kubernetes-based platforms at scale.
• Extensive knowledge of observability tools such as Prometheus, Grafana, Datadog, and OpenTelemetry.
• Strong experience in incident management, including oversight of major incidents.
• Competitive salary and performance-based bonuses.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and growth.
• Flexible working hours and remote work options.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.