Senior Site Reliability Engineer – Fedramp

Posted Sep 17

This is a fully remote position, open to applicants in United States.

📋 Description

• Manage an operational capability for the managed estate, which includes automation, runbooks, and service standards.

• Design observability for regulated cloud environments, focusing on telemetry and log pipelines, service-level objectives, alert quality, and escalation paths.

• Develop and maintain continuous-monitoring evidence pipelines.

• Take ownership of backup and recovery engineering, encompassing tested recovery procedures, measurable recovery objectives, and outage automation.

• Act as the senior escalation point in client environments and handle complex operational incidents.

• Lead incident and problem management, which includes major-event incident command, blameless post-incident reviews, and corrective actions.

• Automate operational tasks using infrastructure-as-code, pipelines, and scripting.

• Collaborate with Engagement Architects and Build teams for transitions into managed operations.

• Maintain on-call responsibility and enhance rotation coverage, alert actionability, and team workload.

• Represent operational posture to clients and assist with renewals and expansions.

• Mentor and guide Site Reliability Engineers and junior personnel.

• Write and peer review code, runbooks, operational design documentation, and compliance artifacts.


⛳️ Requirements

• Bachelor’s degree or higher in a relevant Information Technology field, or an equivalent combination of education and experience.

• Professional or specialty-level certification in AWS, Azure, or GCP; associate-level certification may be acceptable with equivalent proven depth.

• Over 5 years of experience in site reliability engineering, cloud operations, platform engineering, or managed services.

• More than 5 years managing production cloud environments in AWS, Azure, or GCP, including monitoring, incident response, and automation.

• Automation-first mindset with extensive knowledge in Infrastructure-as-Code, CI/CD, scripting, and policy-as-code.

• Strong operational expertise in at least one major cloud platform and familiarity with a second.

• Experience in observability engineering, including metrics, logging, log pipelines, distributed tracing, SLI/SLO definition, and alert design.

• Proven incident response and incident command experience.

• Background in backup, recovery, and resilience engineering.

• Familiarity with NIST 800-53, FedRAMP, or similar security control frameworks.

• Ability to lead technical discussions with clients regarding operational posture, risk, and trade-offs.

• Demonstrated capability to mentor engineers and enhance team productivity.

• Strong communication, organizational, and problem-solving capabilities.

• Proficient documentation skills, including technical diagrams, runbooks, and written descriptions.

• Capability to work independently as well as collaboratively within a team.

• Critical thinking skills to balance security and availability requirements with mission needs.

• Proven experience in owning an operational capability, monitoring platform, or reusable automation utilized by multiple teams or clients.

• Experience as the senior operational escalation point in client-facing managed services, including incident command during major events.

• Advanced knowledge of Infrastructure-as-Code and orchestration/automation tools such as Terraform and Ansible.

• Experience in transitioning environments from build to steady-state operations.


🏝️ Benefits

• Flexible work model enabling employees to choose their work hours and location.

• Paid parental leave.

• Flexible time off policy.

• Reimbursement for certification and training.

• Membership for digital mental health and wellbeing support.

• Comprehensive insurance options available.

• Employee resource groups.

• Opportunities for in-person and virtual events.

• Potential for annual incentive, commission, and/or recognition programs.

People also viewed

Horizon3.ai1 day ago

Staff Site Reliability Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$199.8k – $270k/year
ApplyView job
CLOUD MANTA GmbH1 day ago

Senior DevOps Engineer, Containers & Private Cloud

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)€70k – €80k/year
ApplyView job
Stefanini LATAM1 day ago

Senior DevOps

AR flagArgentina OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Akamai Technologies1 day ago

Principal Site Reliability Engineer – Lead

PL flagPoland OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
PingWind Inc. (SDVOSB)1 day ago

DevSecOps Engineer

US flagAlabama, +1 more stateFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ad Hoc LLC1 day ago

Staff DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$130k – $150k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers