Remotery

Senior Site Reliability Engineer

Posted Jul 28

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the complete operational lifecycle and impact platform architecture, reliability standards, and deployment processes across critical systems.

• Design for reliability and implement automation; support it in a live environment.

• Collaborate on cloud-native infrastructure managing systems that handle millions of provider records.

• Engage in incident response activities, conduct root cause analyses, manage escalation processes, and develop runbooks.

• Construct and maintain Infrastructure as Code, CI/CD pipelines, and operational tools that minimize manual labor and enhance engineering productivity while ensuring reliability.

• Ensure system uptime, mitigate alert fatigue, and establish actionable observability across GKE and Cloud Run without excessive noise.

• Optimize infrastructure scaling, enhance autoscaling performance, resource utilization, and workload efficiency in cloud-native distributed systems.


⛳️ Requirements

• Over 5 years of experience in SRE, DevOps, Platform Engineering, or Infrastructure Engineering — managing production systems at scale where your infrastructure is a critical dependency for others, and failures have significant downstream impacts.

• Proven history of enhancing reliability from end to end: you've troubleshot complex production issues, prevented their recurrence, and established alerting mechanisms to validate it.

• Proficient in Linux systems administration, incident response, and root cause analysis.

• Ability to influence operational standards and guide teams on reliability best practices.

• Extensive hands-on expertise with GCP — GKE, Cloud Run, and containerized workloads at scale.

• Experience in building and maintaining Infrastructure as Code using Terraform and/or Pulumi.

• Familiarity with various deployment strategies and the discernment to determine their appropriate application: rolling deployments, blue/green, canary — along with the rollback procedure for each.

• Knowledge of autoscaling, resource optimization, and infrastructure efficiency for distributed systems.

• Experience in managing infrastructure security, secrets, and access controls in regulated or security-focused environments.

• Strong grasp of Golden Signals monitoring — latency, traffic, errors, saturation — and the ability to turn them into actionable insights instead of noise.

• Experience in designing SLIs, SLOs, error budgets, alerting frameworks, dashboards, and escalation procedures.

• Hands-on experience with observability tools: Google Cloud Monitoring, Datadog, Grafana, Prometheus, or similar platforms.

• A solid understanding of data platform health: aspects like lineage, freshness, and correctness are as important to you as throughput.

• Experience in building and maintaining CI/CD pipelines utilizing GitHub Actions or similar tools.

• Proficiency in scripting or programming languages such as Python, Bash, Go, or similar — you streamline processes through code, not just procedures.

• Experience with Git workflows and contemporary software delivery practices.

• Excellent written and verbal communication skills — you can articulate operational risks to both engineers and product managers in a single discussion.

• Experience operating systems that manage sensitive data or PII in regulated or compliance-adjacent settings.

• Nice to Have:

• - Experience managing large-scale distributed systems or microservices architectures.

• - Familiarity with healthcare, credentialing, or health-tech sectors.

• - Experience utilizing AI-assisted observability or incident response tools.

• - Knowledge of NodeJS, TypeScript, Java, or React application stacks.


🏝️ Benefits

• We prioritize your well-being.

• We offer full coverage of health, dental, and vision insurance premiums for our employees.

• Our US team enjoys unlimited PTO, with a minimum of two weeks off each year to recharge.

• In India, employees receive health insurance, statutory leave benefits, and additional wellness (menstrual) leave for women.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers