
Senior Site Reliability Engineer II
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in North Carolina, +3 more states.
• Oversee the daily operations of the Life Sciences SRE team in alignment with priorities set by the SRE Manager
• Implement reliability and toil-reduction initiatives alongside SREs and contractors
• Collaborate with the Senior Software Architect for Life Sciences to enhance processes, tools, and development strategies
• Engage with teams across Reed Tech to standardize tools, practices, and guidelines
• Shape service level objectives for Life Sciences production systems
• Act as an escalation point during major incidents and direct restoration efforts
• Diagnose and resolve intricate systems and application issues in both development and production settings
• Enhance the SRE framework and contribute to the collective SRE knowledge documentation
• Develop disaster recovery strategies
• Mentor junior SREs and prepare them for on-call responsibilities
• Advocate for AI-assisted development and operational tools
• Create reusable dashboards, establish observability standards, SLOs, and error budgets
• Perform incident analysis, optimize performance, conduct failover testing, manage production recovery, and facilitate blameless post-mortems
• Assist in platform modernization, CI/CD and SDLC enhancements, standard platform adoption, and operational health reporting
• Promote SRE best practices, mentor engineers, identify skill gaps, and lead automation and toil-reduction projects
• Proven experience as a Senior Site Reliability Engineer
• Proficient with Microsoft Azure
• Familiarity with VMs/VMSS, App Service, Azure Functions, and the expanding use of Azure Kubernetes Service (AKS)
• Experience with Azure DevOps Pipelines
• Knowledge of source control and GitHub; GitHub Actions currently under evaluation
• Experience with Azure Monitor and Log Analytics; familiarity with Datadog is a plus
• Proficient in Terraform
• Experience using AI-assisted coding tools such as Claude, GitHub Copilot, or Codex
• Interest or hands-on experience applying AI/ML to observability, automation, or incident response
• Willingness to experiment with and assess emerging AI tools
• Strong understanding of full-stack observability, incident analysis, and performance optimization
• Experience with on-call readiness, mentoring, major incident response, blameless post-mortems, and root-cause analysis
• Advanced knowledge of high-availability systems, resiliency patterns, deployment strategies, and recovery practices
• Experience with failover testing, production recovery, Infrastructure-as-Code, and configuration management tools
• Familiarity with platform modernization, CI/CD, SDLC improvements, and operational health reporting
• Ability to serve as a trusted technical advisor and promote alignment on standards and tools
• Capacity to mentor engineers, identify skill gaps, and advocate for automation and toil reduction
• Flexible working hours
• Shared parental leave
• Study assistance
• Sabbaticals
• Medical, dental and vision benefits
• 401(k) with match
• Employee Share Purchase Plan
• Wellness platform with incentives
• Headspace app subscription
• Employee Assistance and Time-off Programs
• Short-and-Long Term Disability insurance
• Life and Accidental Death Insurance
• Critical Illness insurance
• Hospital Indemnity
• Family bonding and family care leaves
• Adoption and surrogacy benefits
• Health Savings, Health Care, Dependent Care and Commuter Spending Accounts
• Up to two days of paid leave each to engage in Employee Resource Groups and volunteer
Bet On Talent
Virtasant
Ookla
opinov8
Get handpicked remote jobs straight to your inbox weekly.