
Manager, Site Reliability Engineering
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in United States.
• Lead the SRE and DevOps team responsible for supporting Delinea's products.
• Evaluate pull requests and confirm pipeline modifications.
• Optimize monitors and dashboards, execute Datadog queries, and resolve production issues.
• Take ownership of the availability and performance of production environments across Azure and AWS.
• Oversee AKS workloads, ingress and networking, data services, messaging, and CDN/WAF layers.
• Recruit, onboard, mentor, and develop full-time SRE engineers.
• Manage contractor resources, including defining project scopes and assessing deliverables.
• Facilitate one-on-ones, standups, and planning sessions across different time zones.
• Participate in the on-call rotation and act as incident commander for Sev1 and Sev2 incidents.
• Handle detection, triage, mitigation, customer-facing status updates, and post-incident evaluations.
• Ensure that RCAs meet customer-ready criteria and drive preventive measures to completion.
• Enhance observability through SLI/SLO definitions, improvements in alert quality, synthetic coverage, APM instrumentation, log management, and dashboard standards.
• Progress towards supporting Delinea’s FedRAMP High environment.
• Minimize operational toil through automation.
• Report incident metrics, trends, and reliability commitments to leadership while developing improvement plans.
• 6+ years of experience in Site Reliability Engineering, DevOps, or Cloud Operations, with a proven track record of managing production SaaS systems.
• 2+ years of direct experience in people leadership, focusing on performance management, hiring, and coaching.
• Current, hands-on experience with Azure Kubernetes Service, core Azure services (SQL, Redis, Service Bus, Blob Storage), AWS services (SES, EC2, RDS), WAF, Azure DevOps pipelines, Datadog, and Atlassian Jira Service Management.
• Required hands-on experience in both Azure and AWS environments.
• Extensive observability knowledge encompassing metrics, logs, traces, synthetics, SLOs, and alerting strategies.
• Proficient in Datadog or a similar platform, including APM trace analysis and log-based troubleshooting.
• Demonstrated experience serving as incident commander during significant incidents.
• Strong understanding of cloud networking and security fundamentals: load balancing, DNS, TLS and certificate lifecycle, firewalls, VPN, routing, and identity and access management.
• Ability to automate and script using PowerShell, Python, Bash, or similar languages.
• Practical experience with infrastructure-as-code tools such as Terraform, ARM, or Bicep.
• Hands-on experience with multi-region, multi-tenant SaaS architectures, encompassing backup, redundancy, and disaster recovery.
• Excellent written communication skills for customer-facing status reports and incident summaries.
• Availability and willingness to work across time zones and participate in an on-call rotation.
• U.S. work authorization is required; Delinea will not consider candidates who need current or future U.S. work authorization sponsorship.
• Preferred experience operating in a FedRAMP or other regulated environment.
• Preferred familiarity with incident management programs, public status pages, customer notifications, Atlassian Jira Service Management, Confluence, Azure DevOps, and cloud cost optimization.
• Equity.
• Performance-based bonuses or role-based incentive programs.
• Healthcare insurance.
• Pension/retirement matching.
• Comprehensive life insurance.
• Employee assistance program.
• Time off plans.
• Paid company holidays.
• Meaningful work.
• Opportunities for career progression.
• Global team environment.
Bet On Talent
Virtasant
Ookla
opinov8
Get handpicked remote jobs straight to your inbox weekly.