
Senior Site Reliability Engineer – FedRAMP
Posted 10 hours ago

Posted 10 hours ago
This is a fully remote position, open to applicants in United States.
• Take full ownership of the reliability for production SaaS services from start to finish, focusing on availability, performance, and capacity.
• Establish service-level indicators, service-level objectives, and error budgets.
• Develop and optimize monitoring systems in Datadog and Azure Monitor, which includes detection monitors, synthetic checks, dashboards, and alert routing.
• Automate incident responses and convert manual runbook steps into code.
• Engage in on-call rotations and lead the response for high-severity incidents.
• Facilitate incident resolution by coordinating with Support, Engineering, and Product teams.
• Compose post-incident reviews and customer-facing root cause analyses.
• Ensure the completion of preventive actions.
• Address support escalations by reproducing issues, analyzing logs, traces, and network captures, while routing or resolving issues with evidence.
• Construct and maintain infrastructure as code using Terraform and Azure DevOps pipelines.
• Manage the web application firewall, including rule tuning, rate limiting, and addressing false positives.
• Oversee observability platform expenditures, which encompass Datadog indexing, retention, custom metrics, APM, log ingestion, and archived log storage.
• Operate within the FedRAMP High environment, adhering to change control and continuous monitoring processes.
• Enhance the design of on-call rotations, escalation paths, alert quality, runbook coverage, and regional handoffs.
• Collaborate with Support, Security, Product, and Development teams to launch services equipped with monitoring systems, runbooks, and SLOs.
• 8+ years of experience in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product.
• Practical experience with Azure, including AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, along with fundamentals of cloud networking and security.
• Hands-on experience with an observability platform such as Datadog, covering metrics, logs, APM, dashboards, and monitor design.
• Proven track record of managing SLIs, SLOs, and error budgets.
• Experience in incident response, including conducting incident bridges and writing postmortems.
• Familiarity with Kubernetes in production, covering ingress, deployments, resource limits, and troubleshooting failing workloads.
• Experience with infrastructure as code using Terraform.
• Creation and troubleshooting of CI/CD pipelines; familiarity with Azure DevOps is preferred.
• Proficient in scripting with PowerShell and Python.
• Competence in YAML and JSON.
• Solid understanding of networking and web fundamentals: DNS, TLS and certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting.
• Knowledge of redundancy, backup, and disaster recovery strategies in cloud settings.
• Excellent written communication skills.
• Willingness to participate in an on-call rotation, including weekends and emergencies.
• Ability to travel up to 10% of the time.
• Previous government cloud experience is a plus but not mandatory.
• Must not require any form of U.S. work authorization now or in the future.
• Equity.
• Performance-based bonus program or role-specific incentive programs.
• Healthcare insurance.
• Pension/retirement matching.
• Comprehensive life insurance.
• Employee assistance program.
• Time off plans.
• Paid company holidays.
• Opportunities for career progression.
• Engaging work and a culture of innovation.
Bet On Talent
Virtasant
Ookla
opinov8
Get handpicked remote jobs straight to your inbox weekly.