
Senior Site Reliability Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in India.
• Oversee, support, and ensure the reliability, availability, and performance of extensive GeForce NOW production services in both cloud and datacenter settings.
• Engage in production incident triage, troubleshooting, and resolution of intricate infrastructure and application challenges.
• Take part in the on-call rotation, including some weekend duties, to guarantee prompt restoration of services for customers.
• Monitor service health through metrics, logs, traces, and dashboards, proactively identifying reliability, performance, and capacity concerns.
• Collaborate with software engineering, platform, networking, and infrastructure teams to enhance operational readiness, reliability, and service resilience.
• Propel observability initiatives by refining monitoring, alerting, dashboards, and telemetry.
• Develop automation, reduce operational toil, and enhance deployment, recovery, and operational processes.
• Lead and engage in incident responses, root cause analyses, and blameless postmortems.
• Design and create custom tools, automation, and self-service solutions for GeForce NOW.
• Assess operational processes and pinpoint engineering-driven enhancements for reliability, efficiency, and customer satisfaction.
• Contribute to the design, deployment, and management of Kubernetes-based services.
• Bachelor’s degree in Computer Science, Computer Engineering, Information Technology, or a related technical field, or equivalent experience.
• Over 5 years of experience in supporting and operating mission-critical production services in a live-site environment as an SRE, Production Engineer, or a similar position.
• In-depth knowledge of containerization, microservices architecture, and Kubernetes.
• Familiarity with Kubernetes ecosystem components and operational best practices.
• Capability to troubleshoot complex production challenges, determine root causes, and drive issues to resolution.
• Strong understanding of distributed systems across applications, infrastructure, networking, and cloud services.
• Experience with incident management, change management, postmortem reviews, and operational excellence initiatives.
• Practical automation development experience using Python, Go, Bash, or similar programming languages.
• Understanding of SLOs, SLIs, error budgets, KPIs, and service reliability practices.
• Proficient in using Prometheus, Grafana, ELK/OpenSearch, and monitoring and alerting tools.
• Experience operating services in public cloud environments like AWS, Azure, GCP, or equivalent.
• Background in supporting large-scale customer-facing cloud or gaming services.
• Strong operational and troubleshooting expertise in Kubernetes.
• Familiarity with OpenTelemetry.
• Excellent scripting or programming skills with a focus on automation.
• Comprehensive health and wellness benefits.
• Opportunities for professional development and growth.
• Flexible work hours and remote work options.
• Engaging company culture with a focus on collaboration and innovation.
CVS Health
Devoteam
Aspirion
Goodgame Studios
Get handpicked remote jobs straight to your inbox weekly.