
Staff Site Reliability Operations Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in United States, +1 more state.
• Design, optimize, and resolve issues in intricate Layer 1–Layer 7 networking infrastructures.
• Create, scale, and enhance the Grafana Labs observability stack, which includes Grafana, Mimir, Loki, Tempo, and Beyla.
• Implement machine-learning models and automated anomaly detection systems to minimize alert fatigue and forecast bottlenecks.
• Design, scale, secure, and manage production Google Kubernetes Engine clusters.
• Optimize and maintain high-throughput Apache Kafka clusters.
• Guarantee performance, scalability, and disaster recovery preparedness across PostgreSQL, AlloyDB, and BigQuery.
• Integrate AIOps insights into Grafana workflows to facilitate automation of triage, root-cause analysis, and remediation efforts.
• Lead the technical roadmap for distributed infrastructure engineering and GCP cloud-native observability standards.
• Provide mentorship to both senior and junior engineers on advanced debugging, distributed systems, and intelligent operations.
• Over 8 years of experience in SRE, Production Engineering, or roles focused on Distributed Systems infrastructure.
• Demonstrated success and independence in a fully remote engineering environment.
• Extensive technical expertise across OSI Layers 1–7.
• Awareness of physical/fiber infrastructure, switching, BGP, and OSPF.
• Knowledge of TCP congestion control, UDP, and QUIC optimization.
• Understanding of session management, TLS termination, DNS architecture, HTTP/3, and gRPC.
• Expert knowledge of GKE internals, custom controllers, multi-cluster networking, and GitOps workflows.
• Experience in managing high-throughput Apache Kafka pipelines.
• Proven experience handling PostgreSQL, AlloyDB, and BigQuery at scale.
• Practical experience with Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo.
• Experience applying AI/ML techniques for time-series anomaly detection, log clustering, and correlation.
• Advanced expertise in production-scale HashiCorp Terraform for multi-region GCP architectures.
• Strong proficiency in Go and Python programming languages.
• Outstanding written and verbal communication skills.
• In-depth understanding of Google Cloud architecture, Cloud SDN, Cloud Armor, Interconnect, IAM, and cost optimization strategies.
• Familiarity with Linux internals, eBPF-based monitoring, kernel-level networking, Wireshark, and tcpdump.
• Eligibility for bonuses as part of the total compensation package.
• Comprehensive benefits package as referenced by the employer.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.