
Senior Site Reliability Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in United States.
• Take ownership of the complete operational lifecycle and impact platform architecture, reliability standards, and deployment processes across critical systems.
• Design for reliability and implement automation; support it in a live environment.
• Collaborate on cloud-native infrastructure managing systems that handle millions of provider records.
• Engage in incident response activities, conduct root cause analyses, manage escalation processes, and develop runbooks.
• Construct and maintain Infrastructure as Code, CI/CD pipelines, and operational tools that minimize manual labor and enhance engineering productivity while ensuring reliability.
• Ensure system uptime, mitigate alert fatigue, and establish actionable observability across GKE and Cloud Run without excessive noise.
• Optimize infrastructure scaling, enhance autoscaling performance, resource utilization, and workload efficiency in cloud-native distributed systems.
• Over 5 years of experience in SRE, DevOps, Platform Engineering, or Infrastructure Engineering — managing production systems at scale where your infrastructure is a critical dependency for others, and failures have significant downstream impacts.
• Proven history of enhancing reliability from end to end: you've troubleshot complex production issues, prevented their recurrence, and established alerting mechanisms to validate it.
• Proficient in Linux systems administration, incident response, and root cause analysis.
• Ability to influence operational standards and guide teams on reliability best practices.
• Extensive hands-on expertise with GCP — GKE, Cloud Run, and containerized workloads at scale.
• Experience in building and maintaining Infrastructure as Code using Terraform and/or Pulumi.
• Familiarity with various deployment strategies and the discernment to determine their appropriate application: rolling deployments, blue/green, canary — along with the rollback procedure for each.
• Knowledge of autoscaling, resource optimization, and infrastructure efficiency for distributed systems.
• Experience in managing infrastructure security, secrets, and access controls in regulated or security-focused environments.
• Strong grasp of Golden Signals monitoring — latency, traffic, errors, saturation — and the ability to turn them into actionable insights instead of noise.
• Experience in designing SLIs, SLOs, error budgets, alerting frameworks, dashboards, and escalation procedures.
• Hands-on experience with observability tools: Google Cloud Monitoring, Datadog, Grafana, Prometheus, or similar platforms.
• A solid understanding of data platform health: aspects like lineage, freshness, and correctness are as important to you as throughput.
• Experience in building and maintaining CI/CD pipelines utilizing GitHub Actions or similar tools.
• Proficiency in scripting or programming languages such as Python, Bash, Go, or similar — you streamline processes through code, not just procedures.
• Experience with Git workflows and contemporary software delivery practices.
• Excellent written and verbal communication skills — you can articulate operational risks to both engineers and product managers in a single discussion.
• Experience operating systems that manage sensitive data or PII in regulated or compliance-adjacent settings.
• Nice to Have:
• - Experience managing large-scale distributed systems or microservices architectures.
• - Familiarity with healthcare, credentialing, or health-tech sectors.
• - Experience utilizing AI-assisted observability or incident response tools.
• - Knowledge of NodeJS, TypeScript, Java, or React application stacks.
• We prioritize your well-being.
• We offer full coverage of health, dental, and vision insurance premiums for our employees.
• Our US team enjoys unlimited PTO, with a minimum of two weeks off each year to recharge.
• In India, employees receive health insurance, statutory leave benefits, and additional wellness (menstrual) leave for women.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.