Remotery

Staff Site Reliability Operations Engineer

Posted 5 days ago

This is a fully remote position, open to applicants in United States, +1 more state.

📋 Description

• Design, optimize, and resolve issues in intricate Layer 1–Layer 7 networking infrastructures.

• Create, scale, and enhance the Grafana Labs observability stack, which includes Grafana, Mimir, Loki, Tempo, and Beyla.

• Implement machine-learning models and automated anomaly detection systems to minimize alert fatigue and forecast bottlenecks.

• Design, scale, secure, and manage production Google Kubernetes Engine clusters.

• Optimize and maintain high-throughput Apache Kafka clusters.

• Guarantee performance, scalability, and disaster recovery preparedness across PostgreSQL, AlloyDB, and BigQuery.

• Integrate AIOps insights into Grafana workflows to facilitate automation of triage, root-cause analysis, and remediation efforts.

• Lead the technical roadmap for distributed infrastructure engineering and GCP cloud-native observability standards.

• Provide mentorship to both senior and junior engineers on advanced debugging, distributed systems, and intelligent operations.


⛳️ Requirements

• Over 8 years of experience in SRE, Production Engineering, or roles focused on Distributed Systems infrastructure.

• Demonstrated success and independence in a fully remote engineering environment.

• Extensive technical expertise across OSI Layers 1–7.

• Awareness of physical/fiber infrastructure, switching, BGP, and OSPF.

• Knowledge of TCP congestion control, UDP, and QUIC optimization.

• Understanding of session management, TLS termination, DNS architecture, HTTP/3, and gRPC.

• Expert knowledge of GKE internals, custom controllers, multi-cluster networking, and GitOps workflows.

• Experience in managing high-throughput Apache Kafka pipelines.

• Proven experience handling PostgreSQL, AlloyDB, and BigQuery at scale.

• Practical experience with Grafana Enterprise/Cloud, Prometheus/Mimir, Loki, and Tempo.

• Experience applying AI/ML techniques for time-series anomaly detection, log clustering, and correlation.

• Advanced expertise in production-scale HashiCorp Terraform for multi-region GCP architectures.

• Strong proficiency in Go and Python programming languages.

• Outstanding written and verbal communication skills.

• In-depth understanding of Google Cloud architecture, Cloud SDN, Cloud Armor, Interconnect, IAM, and cost optimization strategies.

• Familiarity with Linux internals, eBPF-based monitoring, kernel-level networking, Wireshark, and tcpdump.


🏝️ Benefits

• Eligibility for bonuses as part of the total compensation package.

• Comprehensive benefits package as referenced by the employer.

People also viewed

CVS Health5 hours ago

Salesforce DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$83.4k – $166.9k/year
ApplyView job
Devoteam6 hours ago

Data, AWS DevSecOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Aspirion6 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Goodgame Studios7 hours ago

Senior Agentic Engineer – Java Backend, DevOps

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Instacart7 hours ago

Site Reliability Engineer II

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$133k – $169k/year
ApplyView job
Logicalis Spain8 hours ago

DevOps Engineer

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€40k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers