
Senior Site Reliability Engineer – Fedramp
Posted Sep 17

Posted Sep 17
This is a fully remote position, open to applicants in United States.
• Manage an operational capability for the managed estate, which includes automation, runbooks, and service standards.
• Design observability for regulated cloud environments, focusing on telemetry and log pipelines, service-level objectives, alert quality, and escalation paths.
• Develop and maintain continuous-monitoring evidence pipelines.
• Take ownership of backup and recovery engineering, encompassing tested recovery procedures, measurable recovery objectives, and outage automation.
• Act as the senior escalation point in client environments and handle complex operational incidents.
• Lead incident and problem management, which includes major-event incident command, blameless post-incident reviews, and corrective actions.
• Automate operational tasks using infrastructure-as-code, pipelines, and scripting.
• Collaborate with Engagement Architects and Build teams for transitions into managed operations.
• Maintain on-call responsibility and enhance rotation coverage, alert actionability, and team workload.
• Represent operational posture to clients and assist with renewals and expansions.
• Mentor and guide Site Reliability Engineers and junior personnel.
• Write and peer review code, runbooks, operational design documentation, and compliance artifacts.
• Bachelor’s degree or higher in a relevant Information Technology field, or an equivalent combination of education and experience.
• Professional or specialty-level certification in AWS, Azure, or GCP; associate-level certification may be acceptable with equivalent proven depth.
• Over 5 years of experience in site reliability engineering, cloud operations, platform engineering, or managed services.
• More than 5 years managing production cloud environments in AWS, Azure, or GCP, including monitoring, incident response, and automation.
• Automation-first mindset with extensive knowledge in Infrastructure-as-Code, CI/CD, scripting, and policy-as-code.
• Strong operational expertise in at least one major cloud platform and familiarity with a second.
• Experience in observability engineering, including metrics, logging, log pipelines, distributed tracing, SLI/SLO definition, and alert design.
• Proven incident response and incident command experience.
• Background in backup, recovery, and resilience engineering.
• Familiarity with NIST 800-53, FedRAMP, or similar security control frameworks.
• Ability to lead technical discussions with clients regarding operational posture, risk, and trade-offs.
• Demonstrated capability to mentor engineers and enhance team productivity.
• Strong communication, organizational, and problem-solving capabilities.
• Proficient documentation skills, including technical diagrams, runbooks, and written descriptions.
• Capability to work independently as well as collaboratively within a team.
• Critical thinking skills to balance security and availability requirements with mission needs.
• Proven experience in owning an operational capability, monitoring platform, or reusable automation utilized by multiple teams or clients.
• Experience as the senior operational escalation point in client-facing managed services, including incident command during major events.
• Advanced knowledge of Infrastructure-as-Code and orchestration/automation tools such as Terraform and Ansible.
• Experience in transitioning environments from build to steady-state operations.
• Flexible work model enabling employees to choose their work hours and location.
• Paid parental leave.
• Flexible time off policy.
• Reimbursement for certification and training.
• Membership for digital mental health and wellbeing support.
• Comprehensive insurance options available.
• Employee resource groups.
• Opportunities for in-person and virtual events.
• Potential for annual incentive, commission, and/or recognition programs.
Horizon3.ai
CLOUD MANTA GmbH
Stefanini LATAM
Akamai Technologies
Get handpicked remote jobs straight to your inbox weekly.