
Site Reliability Engineering Manager
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Florida.
• Lead, mentor, and cultivate a US-based team of Site Reliability Engineers.
• Conduct regular one-on-one meetings, performance evaluations, and discussions on career development.
• Oversee hiring, onboarding, and retention initiatives as the team expands.
• Promote a culture of ownership, blameless postmortems, and ongoing improvement.
• Manage day-to-day production operations, ensuring timely incident triage, resolution, and escalation.
• Act as an escalation point and incident commander for significant production incidents.
• Facilitate problem management and root cause analysis processes.
• Carry PagerDuty on-call escalation responsibilities for urgent issues.
• Monitor and report on operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
• Enhance system reliability, observability, and resilience using Datadog and related tools.
• Promote automation, self-healing capabilities, and the maturity of runbooks.
• Collaborate with Development, DevOps, DevSecOps, and Engineering teams to integrate reliability into the software development life cycle (SDLC).
• Contribute hands-on to tooling, automation, and technical assessments as required.
• Work closely with SRE leadership in India to ensure effective follow-the-sun coverage.
• Represent the US SRE organization in cross-functional planning and operational reviews.
• Communicate effectively with both technical and non-technical stakeholders.
• Maintain high-quality documentation for incidents, postmortems, runbooks, and operational procedures.
• Ensure compliance with healthcare and fintech standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
• 5–8 years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
• 1–2+ years of experience in leading, mentoring, or managing engineers.
• Proven success in a player-coach leadership model.
• Strong hands-on experience with production incident management and escalation procedures.
• Proficiency with Datadog or similar observability tools.
• Practical experience with Kubernetes and Docker in production settings.
• Strong scripting or programming skills in PowerShell, Bash, Python, Java, or C#.
• Experience with Helm, CI/CD pipelines, and deployment automation.
• Familiarity with ITIL processes and Agile methodologies.
• Experience working with SQL, MySQL, or NoSQL databases.
• Excellent communication and stakeholder management capabilities.
• Willingness to participate in PagerDuty on-call escalation and work within a global follow-the-sun operating framework.
• Competitive salary and comprehensive benefits package.
• Unlimited paid time off (PTO).
• Fully remote work environment (for US-based employees).
• Opportunity to lead and develop a high-impact Site Reliability Engineering organization.
• Exposure to modern cloud-native technologies and large-scale reliability challenges.
• Collaborative culture that emphasizes innovation, learning, and continuous improvement.
• Meaningful work that directly contributes to healthcare technology and affects millions of members.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.