
Site Reliability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in India.
• Design, implement, and uphold the reliability, availability, and performance of essential JumpCloud systems and APIs across AWS and GCP.
• Operationalize SLIs, SLOs, and error budgets in collaboration with core application teams.
• Construct and enhance end-to-end observability across microservices and cloud infrastructure utilizing tools like Datadog.
• Establish monitoring based on Golden Signals: latency, traffic, errors, and saturation.
• Engage in on-call rotations, incident response, and conduct blameless post-incident reviews.
• Oversee production Kubernetes EKS clusters using GitOps workflows such as Argo CD and Kargo.
• Provision and secure multi-cloud infrastructure through modular Terraform.
• Create and maintain disaster recovery dashboards, runbooks, multi-region failover automation, and validation tests in line with RTO/RPO targets.
• Develop production-quality Python or Go scripts and automation tools to reduce operational toil.
• Leverage AI-assisted development tools such as Cursor, Claude Code, and GitHub Copilot for scripting, runbook generation, and incident triage.
• A minimum of 5 years of professional software engineering experience in SRE, DevOps, or Platform Engineering managing 24/7 mission-critical systems.
• Proficiency in Python or Go for developing SRE tools, custom automation, and cloud integrations.
• Hands-on experience with Kubernetes, container orchestration, and GitOps pipelines like Argo CD.
• Experience in writing, maintaining, and modularizing Terraform configurations.
• Familiarity with operating AWS workloads, including EKS, IAM, VPC networking, Route53, and ALB/NLB, or GCP workloads.
• Practical knowledge of FinOps, cost-allocation tagging, resource right-sizing, and cloud-spend dashboards.
• Experience in building disaster recovery dashboards, conducting failover drills, and setting up monitoring for system health and recovery metrics.
• Proficient with Datadog or similar tools, PagerDuty, alerting hygiene, and SLI/SLO frameworks.
• Operational experience in configuring and troubleshooting production service meshes such as Istio and high-availability proxy solutions like HAProxy or NGINX.
• Strong problem-solving abilities and a proven record of enhancing operational efficiency through coding.
• Excellent team player with alignment to company core values.
• Willingness and capability to participate in on-call shifts.
• Fluent in spoken and written English.
• Preferred: experience with GitHub Actions or GitLab Pipelines.
• Preferred: basic understanding of chaos engineering or resilience testing.
• Preferred: familiarity with HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator.
• Preferred: basic knowledge of DevSecOps tools and infrastructure-as-code vulnerability remediation.
• Remote-first work environment within India.
• Opportunity to thrive in a fast-paced, SaaS-based setting.
• Professional development and expertise-sharing opportunities.
• Collaboration with skilled global teams.
• Employee input in product and feature development.
• Supportive leadership team and board.
• Commitment to equal opportunity employment.
• No third-party resumes will be accepted.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.