Senior Site Reliability Engineer – FedRAMP

Posted 10 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take full ownership of the reliability for production SaaS services from start to finish, focusing on availability, performance, and capacity.

• Establish service-level indicators, service-level objectives, and error budgets.

• Develop and optimize monitoring systems in Datadog and Azure Monitor, which includes detection monitors, synthetic checks, dashboards, and alert routing.

• Automate incident responses and convert manual runbook steps into code.

• Engage in on-call rotations and lead the response for high-severity incidents.

• Facilitate incident resolution by coordinating with Support, Engineering, and Product teams.

• Compose post-incident reviews and customer-facing root cause analyses.

• Ensure the completion of preventive actions.

• Address support escalations by reproducing issues, analyzing logs, traces, and network captures, while routing or resolving issues with evidence.

• Construct and maintain infrastructure as code using Terraform and Azure DevOps pipelines.

• Manage the web application firewall, including rule tuning, rate limiting, and addressing false positives.

• Oversee observability platform expenditures, which encompass Datadog indexing, retention, custom metrics, APM, log ingestion, and archived log storage.

• Operate within the FedRAMP High environment, adhering to change control and continuous monitoring processes.

• Enhance the design of on-call rotations, escalation paths, alert quality, runbook coverage, and regional handoffs.

• Collaborate with Support, Security, Product, and Development teams to launch services equipped with monitoring systems, runbooks, and SLOs.


⛳️ Requirements

• 8+ years of experience in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product.

• Practical experience with Azure, including AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, along with fundamentals of cloud networking and security.

• Hands-on experience with an observability platform such as Datadog, covering metrics, logs, APM, dashboards, and monitor design.

• Proven track record of managing SLIs, SLOs, and error budgets.

• Experience in incident response, including conducting incident bridges and writing postmortems.

• Familiarity with Kubernetes in production, covering ingress, deployments, resource limits, and troubleshooting failing workloads.

• Experience with infrastructure as code using Terraform.

• Creation and troubleshooting of CI/CD pipelines; familiarity with Azure DevOps is preferred.

• Proficient in scripting with PowerShell and Python.

• Competence in YAML and JSON.

• Solid understanding of networking and web fundamentals: DNS, TLS and certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting.

• Knowledge of redundancy, backup, and disaster recovery strategies in cloud settings.

• Excellent written communication skills.

• Willingness to participate in an on-call rotation, including weekends and emergencies.

• Ability to travel up to 10% of the time.

• Previous government cloud experience is a plus but not mandatory.

• Must not require any form of U.S. work authorization now or in the future.


🏝️ Benefits

• Equity.

• Performance-based bonus program or role-specific incentive programs.

• Healthcare insurance.

• Pension/retirement matching.

• Comprehensive life insurance.

• Employee assistance program.

• Time off plans.

• Paid company holidays.

• Opportunities for career progression.

• Engaging work and a culture of innovation.

People also viewed

Bet On Talent6 hours ago

Senior DevOps Engineer

EuropeFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Virtasant6 hours ago

Build & Release Support Engineer – CI/CD

MX flagMexico, +5 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ookla6 hours ago

Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$90k – $100k/year
ApplyView job
opinov87 hours ago

Senior DevOps Engineer, Media and Advertising Industry

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
GFT Technologies7 hours ago

DevOps Specialist

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
RELX10 hours ago

Senior Site Reliability Engineer II

US flagNorth Carolina, +3 more statesFull-timeDevOps & Site Reliability Engineer (SRE)$104.9k – $174.7k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers