Site Reliability Engineer

Posted 2 days ago

This is a fully remote position, open to applicants in India.

📋 Description

• Design, implement, and uphold the reliability, availability, and performance of essential JumpCloud systems and APIs across AWS and GCP.

• Operationalize SLIs, SLOs, and error budgets in collaboration with core application teams.

• Construct and enhance end-to-end observability across microservices and cloud infrastructure utilizing tools like Datadog.

• Establish monitoring based on Golden Signals: latency, traffic, errors, and saturation.

• Engage in on-call rotations, incident response, and conduct blameless post-incident reviews.

• Oversee production Kubernetes EKS clusters using GitOps workflows such as Argo CD and Kargo.

• Provision and secure multi-cloud infrastructure through modular Terraform.

• Create and maintain disaster recovery dashboards, runbooks, multi-region failover automation, and validation tests in line with RTO/RPO targets.

• Develop production-quality Python or Go scripts and automation tools to reduce operational toil.

• Leverage AI-assisted development tools such as Cursor, Claude Code, and GitHub Copilot for scripting, runbook generation, and incident triage.


⛳️ Requirements

• A minimum of 5 years of professional software engineering experience in SRE, DevOps, or Platform Engineering managing 24/7 mission-critical systems.

• Proficiency in Python or Go for developing SRE tools, custom automation, and cloud integrations.

• Hands-on experience with Kubernetes, container orchestration, and GitOps pipelines like Argo CD.

• Experience in writing, maintaining, and modularizing Terraform configurations.

• Familiarity with operating AWS workloads, including EKS, IAM, VPC networking, Route53, and ALB/NLB, or GCP workloads.

• Practical knowledge of FinOps, cost-allocation tagging, resource right-sizing, and cloud-spend dashboards.

• Experience in building disaster recovery dashboards, conducting failover drills, and setting up monitoring for system health and recovery metrics.

• Proficient with Datadog or similar tools, PagerDuty, alerting hygiene, and SLI/SLO frameworks.

• Operational experience in configuring and troubleshooting production service meshes such as Istio and high-availability proxy solutions like HAProxy or NGINX.

• Strong problem-solving abilities and a proven record of enhancing operational efficiency through coding.

• Excellent team player with alignment to company core values.

• Willingness and capability to participate in on-call shifts.

• Fluent in spoken and written English.

• Preferred: experience with GitHub Actions or GitLab Pipelines.

• Preferred: basic understanding of chaos engineering or resilience testing.

• Preferred: familiarity with HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator.

• Preferred: basic knowledge of DevSecOps tools and infrastructure-as-code vulnerability remediation.


🏝️ Benefits

• Remote-first work environment within India.

• Opportunity to thrive in a fast-paced, SaaS-based setting.

• Professional development and expertise-sharing opportunities.

• Collaboration with skilled global teams.

• Employee input in product and feature development.

• Supportive leadership team and board.

• Commitment to equal opportunity employment.

• No third-party resumes will be accepted.

People also viewed

Koniag Government Services1 day ago

Architect/DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
FP Markets (First Prudential Markets)1 day ago

Senior DevOps Engineer

AM flagArmenia, +4 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Modern Campus2 days ago

Senior DevOps Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
InRule2 days ago

Site Reliability Engineer

US flagUnited States, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

US flagUnited States, +38 more locationsFull-timeDevOps & Site Reliability Engineer (SRE)$179.4k – $272.8k/year
ApplyView job
Thumbtack2 days ago

Senior Software Engineer, Site Reliability Engineering

CA flagCanada, +1 more countryFull-timeDevOps & Site Reliability Engineer (SRE)C$180.2k – C$233.2k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers