
Senior Site Reliability Engineer
Posted Jul 17

Posted Jul 17
This is a fully remote position, open to applicants in India.
• Engage in on-call rotations as the lead technical authority. Serve as the Incident Commander during high-severity incidents by initiating war rooms, coordinating multi-disciplinary teams, and delivering clear status updates.
• Implement code to reveal high-cardinality metrics and distributed traces. Work collaboratively to define, measure, and uphold Service Level Objectives (SLOs) and Error Budgets alongside product owners.
• Develop high-quality, production-ready code (in Java, Go, or Python) to create internal tools, automation platforms, and self-healing mechanisms that remove the need for manual operator intervention.
• Collaborate with Product Engineering teams during the design phase to ensure that new services are constructed with reliability, scalability, and observability patterns (circuit breakers, rate limiting, backpressure, fallback strategies) from the outset.
• Evaluate system performance and traffic patterns to forecast future capacity requirements. Perform load testing and chaos engineering experiments to validate system resilience under failure scenarios.
• A minimum of 5 - 7 years of experience in Site Reliability Engineering (SRE) or Backend Engineering with a proven ability to write clean, efficient, and well-tested code in Java, Go, Rust, or Python.
• A thorough understanding of distributed systems architecture and design patterns. You have a strong grasp of microservices fundamentals, event-driven architectures, and the core principles necessary for building scalable systems.
• Significant experience with Google Cloud Platform (GCP) or other similar cloud services (AWS/Azure). You are skilled in managing production workloads on Kubernetes (GKE/EKS) and troubleshooting cluster/infrastructure challenges.
• Proven experience in designing observability strategies utilizing OpenTelemetry, Prometheus, New Relic, Datadog, or SigNoz to enhance system visibility.
• Familiarity with operating and optimizing production data stores (e.g., PostgreSQL, MySQL) and streaming platforms (e.g., Kafka, RabbitMQ) in a high-throughput setting.
• Culture - We prioritize our people and their well-being. We’ve fostered an environment where every opinion is valued, and every voice is heard. We respect and support one another, emphasizing our humanity above all.
• Learning - Our environment is focused on learning and development, with a strong emphasis on knowledge sharing, training, and regular internal technical discussions.
• Compensation - You will receive a competitive salary, pension, health insurance, annual bonus, and additional benefits.
The Codest
IRIUM
Sólides
Resilinc
Get handpicked remote jobs straight to your inbox weekly.