
Senior Site Reliability Engineer
Posted Jun 23

Posted Jun 23
This is a fully remote position, open to applicants in India.
• Take charge of the reliability, scalability, and functionality of the SigNoz cloud platform.
• Ensure that a petabyte-scale observability system remains quick and reliable.
• Enhance the ingest path to withstand bursts while preserving data freshness.
• Manage and optimize ClickHouse and the data layer for both performance and cost efficiency.
• Oversee Kubernetes infrastructure, including cluster operations, upgrades, and multi-tenancy.
• Contribute to achieving world-class observability for SigNoz.
• Collaborate with a talented team on various tasks, including SLOs/SLIs, incident response, and tooling.
• 5–8 years of experience in SRE, infrastructure, or platform/backend roles managing production systems at scale.
• Extensive, hands-on experience with Kubernetes.
• Strong understanding of distributed systems failure modes, performance debugging, and capacity planning.
• Proficient in coding (Go preferred).
• Passionate about open source, ideally with previous contributions to OSS projects.
• Comfortable working in a high-ownership, fast-paced, remote-first environment.
• Excellent communication skills—able to write clear runbooks and technical documentation and articulate trade-offs effectively.
• Remote-first, async-friendly culture.
Ontrac Solutions
CyberSheath
Ontrac Solutions
NVIDIA
Get handpicked remote jobs straight to your inbox weekly.