
Manager, Site Reliability Engineering
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Oversee the Platform SRE and DevOps team that supports Delinea products.
• Assess pull requests and confirm pipeline modifications.
• Optimize monitors and dashboards, execute Datadog queries, and resolve production challenges.
• Take responsibility for the availability and performance of production environments across Azure and AWS.
• Manage AKS workloads, ingress and networking, data services, messaging, and CDN/WAF layers.
• Recruit, onboard, mentor, and develop full-time SRE engineers.
• Direct and oversee contractor resources, including defining project scope and reviewing deliverables.
• Conduct one-on-ones, standups, and planning meetings across various time zones.
• Engage in the on-call rotation and act as incident commander for Sev1 and Sev2 incidents.
• Manage detection, triage, mitigation, customer-facing status communications, and post-incident evaluations.
• Ensure that RCAs are up to customer-ready standards and drive preventive measures to completion.
• Enhance observability through SLI/SLO definitions, improvements in alert quality, synthetic coverage, APM instrumentation, log hygiene, and dashboard standards.
• Progress into supporting Delinea’s FedRAMP High environment.
• Minimize operational toil through automation.
• Report incident metrics, trends, and reliability commitments to leadership and formulate improvement plans.
• 6+ years of experience in Site Reliability Engineering, DevOps, or Cloud Operations, demonstrating ownership of production SaaS systems.
• 2+ years in direct people leadership roles, including performance management, recruitment, and coaching.
• Current, hands-on production experience with Azure Kubernetes Service, essential Azure services (SQL, Redis, Service Bus, Blob Storage), AWS services (SES, EC2, RDS), WAF, Azure DevOps pipelines, Datadog, and Atlassian Jira Service Management.
• Required hands-on experience with both Azure and AWS.
• Extensive observability knowledge encompassing metrics, logs, traces, synthetics, SLOs, and alerting strategies.
• Proficient in using Datadog or a similar platform, including APM trace analysis and log-based troubleshooting.
• Documented experience serving as incident commander during major incidents.
• Strong understanding of cloud networking and security principles: load balancing, DNS, TLS and certificate lifecycle, firewalls, VPN, routing, and identity and access management.
• Capability in automation and scripting with PowerShell, Python, Bash, or similar languages.
• Practical experience in infrastructure-as-code with Terraform, ARM, or Bicep.
• Experience with multi-region, multi-tenant SaaS architectures, including backup, redundancy, and disaster recovery.
• Excellent written communication skills for customer-facing status updates and incident reports.
• Willingness and availability to work across time zones and take part in an on-call rotation.
• U.S. work authorization is required; Delinea will not consider candidates requiring current or future U.S. work authorization sponsorship.
• Direct experience in a FedRAMP or other regulated environment is preferred.
• Familiarity with incident management programs, public status pages, customer notifications, Atlassian Jira Service Management, Confluence, Azure DevOps, and cloud cost optimization is preferred.
• Equity
• Performance-based bonus or role-based incentive programs
• Healthcare insurance
• Pension/retirement matching
• Comprehensive life insurance
• Employee assistance program
• Time off plans
• Paid company holidays
• Meaningful work
• Career progression
• Global team environment
FCamara Consulting & Training
Sequoia Connect
FundCount
GE Vernova
Get handpicked remote jobs straight to your inbox weekly.