
Site Reliability Engineering Lead
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Lead and mentor a group of Site Reliability Engineers.
• Develop and implement the SRE strategy, standards, best practices, and operational frameworks.
• Enhance application reliability, scalability, security, performance, and resilience in collaboration with product and platform teams.
• Establish and uphold Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
• Oversee major incident management, root cause analysis, problem management, and conduct post-incident reviews.
• Propel cloud modernization and application migrations to Azure, AWS, and containerized environments.
• Advocate for automation and Infrastructure as Code methodologies utilizing Terraform, GitHub, GitLab, Jenkins, and Ansible.
• Formulate observability strategies incorporating metrics, logs, traces, alerting, and dashboards.
• Collaborate with security and compliance teams to ensure adherence to enterprise security and regulatory requirements.
• Lead architecture reviews and provide guidance on cloud-native and resilient application design.
• Manage capacity planning, performance optimization, cost management, and operational efficiency.
• Establish engineering guardrails, governance controls, and standards for production deployment.
• Facilitate the shift towards DevOps and SRE practices.
• Manage operational risk and ensure business continuity and disaster recovery preparedness.
• Prioritize reliability enhancements and platform investments in consultation with stakeholders.
• Conduct resource planning and assist with hiring, onboarding, and career development initiatives.
• Set team objectives that align with both business and technology strategies.
• Serve as a subject matter expert in reliability engineering and a trusted advisor.
• Over 8 years of experience in Cloud Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or Site Reliability Engineering.
• At least 2 years of leadership or people management experience overseeing engineering teams.
• Bachelor’s degree in Computer Science, Engineering, Information Systems, or equivalent practical experience.
• Preferably possess Azure, AWS, Kubernetes, Terraform, or related certifications.
• Demonstrated success in leading reliability and operational excellence initiatives in large-scale enterprise settings.
• Strong leadership skills with experience managing technical engineering teams.
• In-depth knowledge of Site Reliability Engineering, DevOps, Cloud Engineering, or Platform Engineering.
• Extensive experience with Azure and/or AWS.
• Solid understanding of Kubernetes, AKS, EKS, containerization, Docker, and cloud-native architectures.
• Proficient in Terraform and Ansible.
• Strong background in observability platforms such as Grafana, Prometheus, OpenTelemetry, Splunk, Dynatrace, or Datadog.
• Experience managing large-scale production environments with high availability demands.
• Knowledge of security, compliance, networking, and cloud governance principles.
• Experience in designing highly available, fault-tolerant, and resilient systems.
• Proficient in programming languages such as Python, Go, PowerShell, Bash, or C#.
• Familiarity with CI/CD pipelines and software delivery automation.
• Exceptional troubleshooting and problem-solving abilities.
• Excellent communication and stakeholder management skills.
• Strong documentation and presentation capabilities.
• Ability to influence technical direction across multiple engineering teams.
• Eligibility for annual incentive bonuses.
• Country-specific employee benefits.
• Support for accommodation or adjustments during the hiring process.
VALCE Talent Solutions
4Pharma Ltd
MTP Brasil
BlackSky
Get handpicked remote jobs straight to your inbox weekly.