
System Reliability Engineering Lead
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Act as the hands-on technical expert for ensuring production stability across the GridOS SaaS portfolio.
• Manage Change Management processes and make decisions to approve or halt production deployments based on system health.
• Promote high reliability and engineering excellence within a distributed team environment.
• Design and implement standardized, secure cloud infrastructure provisioning.
• Automate account provisioning to expedite customer onboarding.
• Define and develop the standardized Middle-Mile software delivery platform utilizing Backstage, ArgoCD, and GitHub Actions.
• Create global handover protocols and ensure 24/7 operational coverage across US, India, and Mexico time zones.
• Establish and manage enterprise-wide SLOs, SLIs, and error budgets.
• Serve as the ultimate technical authority for production releases, enforcing security and performance quality gates.
• Implement Canary and Blue/Green deployment strategies with automated rollback functionalities.
• Develop and enhance the SRE Center for Enablement, providing coaching, templates, and reliability patterns.
• Lead incident response efforts for Sev1/Sev2 events and P1 escalations.
• Facilitate blameless Root Cause Analysis and oversee the post-incident lifecycle.
• Architect and validate backup and disaster recovery strategies, including cross-region failover and automated recovery testing.
• Manage FinOps, optimize cloud costs, and engage in long-term capacity planning.
• Act as the primary SRE contact for North American utility clients.
• Engage in customer reviews, incident communications, and service health reporting.
• Lead a distributed team of 8 SRE engineers located in Hyderabad and Querétaro.
• Establish technical direction, assign tasks, oversee deliverables, mentor engineers, and provide performance feedback to the relevant people leader.
• Travel up to 10% to customer sites and team locations as required.
• Extensive knowledge of AWS core services: EC2, EKS, RDS, S3, and IAM.
• Familiarity with AWS management tools such as CloudTrail and CloudWatch.
• Advanced understanding of Kubernetes internals and EKS cluster operations across multi-region architectures.
• Expert-level proficiency in ArgoCD, GitHub Actions, and GitOps-first workflows.
• Strong skills in Infrastructure as Code using Terraform.
• Proficiency in configuration management with Ansible.
• Practical experience with Prometheus, Grafana, Splunk or Datadog, and OpenTelemetry.
• Experience in cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms.
• Over 12 years of experience in software engineering, cloud operations, or infrastructure roles.
• 8-10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications.
• Proven track record of leading distributed engineering teams in a player-coach capacity while remaining directly involved with architecture, automation, and incident response.
• Exceptional troubleshooting abilities under pressure and a proactive mindset towards investigation and inspection.
• Experience collaborating directly with enterprise customers on production reliability, incident communication, and service-level reporting.
• Must pass customer-mandated background screening for access to critical infrastructure environments.
• Must be legally authorized to work in the United States.
• Must complete a drug screen, as applicable.
• General shift during US business hours and on-call availability for P1/Sev1 incidents.
• Up to 10% travel to customer sites and team locations.
• Desired: Knowledge or experience with NERC CIP, SOC2, ISO 27001, or IEC 62443.
• Desired: Experience in highly regulated sectors, such as utilities, financial services, or critical national infrastructure.
• Desired certifications: AWS DevOps Engineer—Professional or Solutions Architect—Associate/Professional, CKA, SRE Practitioner, and AWS FinOps Practitioner or equivalent.
• Discretionary annual bonus.
• Coverage for medical, dental, vision, and prescription drugs.
• Access to a Health Coach from GE Vernova, a resource available 24/7.
• Employee Assistance Program offering 24/7 confidential assessments, counseling, and referral services.
• GE Vernova Retirement Savings Plan.
• Tax-advantaged 401(k) savings plan with company matching contributions and retirement contributions.
• Resources and financial planning consultants available through Fidelity.
• Tuition assistance.
• Adoption assistance.
• Paid parental leave.
• Disability benefits.
• Life insurance coverage.
• 12 paid holidays.
• Permissive time off policy.
• Opportunities for professional development.
• Relocation assistance not provided.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.