
Senior Site Reliability Engineer
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in India.
• Design, scale, and continually enhance the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication systems on AWS/GCP.
• Engineer, establish, and sustain Disaster Recovery procedures, multi-region failover automation, and business continuity plans.
• Spearhead the design and implementation of SLIs, SLOs, and Error Budget frameworks.
• Propel an end-to-end observability strategy utilizing Datadog and Golden Signals monitoring.
• Oversee on-call escalation and major incident management while ensuring compliance with 99.99% availability SLAs.
• Conduct blameless post-incident reviews and execute systemic root-cause remediation actions.
• Design, manage, and scale production Kubernetes EKS clusters employing GitOps workflows with Argo CD and Kargo.
• Create and uphold Infrastructure-as-Code using Terraform across multi-account, multi-region cloud setups.
• Develop FinOps and cost-optimization dashboards for multi-cloud expenditures, unit economics, and resource utilization.
• Build production-grade tooling in Python or Go, including platform automation and custom integrations.
• Advocate for AI-assisted software development workflows utilizing Cursor, Claude Code, and GitHub Copilot.
• Write operational runbooks and architecture decision records.
• Mentor mid-level and junior engineers.
• 8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering managing 24/7 mission-critical, highly available distributed systems.
• Bachelor's degree in Computer Science, Software Engineering, or a related technical field.
• Advanced proficiency in Python or Go.
• Practical experience with production EKS/GKE cluster lifecycles, ingress/egress management, networking, RBAC, and GitOps tools like Argo CD.
• Strong Terraform expertise across complex multi-account AWS environments, including IAM, VPCs, Transit Gateway, ALB/NLB, and Route53.
• Proven experience in driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and FinOps dashboards.
• Experience in designing and testing multi-region Disaster Recovery architectures and automating failover systems.
• Competence in defining SLIs/SLOs, managing PagerDuty schedules, and enhancing production observability platforms.
• Familiarity with designing and operating enterprise service meshes like Istio or Linkerd and production ingress/proxy systems such as HAProxy or NGINX.
• Capability to lead technical discussions, draft architectural design documents/RFCs, and mentor engineering colleagues.
• Excellent problem-solving, communication, and collaboration abilities.
• Preferred: basic understanding of chaos engineering; secrets management architectures; DevSecOps practices; identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions.
• Must reside in and be authorized to work in India.
• Proficient spoken and written English is required.
• Remote work options available within India.
• Collaborate with talent from over 15 countries.
• Opportunities for professional growth and knowledge sharing.
• Supportive and collaborative workplace culture.
• Equal opportunity employer.
Koniag Government Services
FP Markets (First Prudential Markets)
Modern Campus
InRule
Get handpicked remote jobs straight to your inbox weekly.