Senior Site Reliability Engineer

Posted 6 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take charge of the reliability of production systems and establish how reliability is evaluated.

• Collaborate with stakeholders to define Service Level Objectives (SLOs) and Service Level Indicators (SLIs), create error budgets, and utilize them for guiding release decisions.

• Manage and enhance the SaaS monitoring, alerting, and logging infrastructure to fulfill compliance standards and proactively identify issues.

• Engage in on-call production incident response and assist engineers in resolving customer-related challenges.

• Facilitate blameless postmortems, pinpoint root causes, and implement sustainable solutions.

• Identify, assess, and automate repetitive manual engineering tasks.

• Ensure that automation and documentation remain intact through changes in ownership.

• Oversee the design, construction, and upkeep of essential infrastructure for secure and efficient product delivery.

• Prioritize reliability and platform tasks, secure managerial approval, and adjust according to business requirements.

• Evaluate risk versus impact for high-profile systems and implement measurable, reversible improvements.

• Make determinations on build versus buy and suggest established tools.

• Collaborate with engineering teams on Continuous Integration/Continuous Deployment (CI/CD), testing, canary releases, and automated rollback strategies.

• Work alongside security teams to maintain compliance in infrastructure and operations while proactively identifying risks.

• Mentor engineers on reliability, Azure, AWS, Terraform, and operational methodologies.

• Develop documentation for system operations and lead solutions, clearly outlining trade-offs.


⛳️ Requirements

• 8+ years of experience in site reliability, DevOps, or infrastructure engineering, with complete ownership of production systems.

• Experience in building and managing AWS and Azure accounts and resources in accordance with the Well-Architected Framework.

• Familiarity with AWS services such as ECS, Fargate, RDS, Lambda, SNS, SQS, S3, EventBridge, and Step Functions.

• Proven experience in deploying Docker-based applications to production and understanding of reliable container management.

• Hands-on proficiency with Azure, AWS, and traditional data centers, including the management of Windows VMs and IIS.

• In-depth knowledge of SQL and relational database administration, including query optimization, index tuning, and resolving high-load production issues; SQL Server experience preferred.

• Experience with Infrastructure-as-Code practices using Terraform.

• Practical familiarity with DevOps principles, the 12-Factor App methodology, least privilege access, and zero-trust architecture.

• Experience leading cross-departmental requirements and defining Service Level Agreements (SLAs).

• Proven track record in defining and operating against SLOs, as well as leading incident response and postmortems for production outages.

• Knowledge of ITIL-aligned service management, covering incident, problem, and change management.

• Experience in mentoring engineers on reliability and delivery best practices.

• Must possess United States Citizenship.

• Ability to meet security investigation and eligibility criteria for access to classified (Public Trust) information.

• Preferred certifications in Azure or AWS, such as AZ-104, AZ-305, or AWS Certified Solutions Architect.

• Experience with Azure Government or other federal cloud platforms preferred.

• Background in supporting the VA, DoD, or other federal health initiatives, including FedRAMP or ATO-related work preferred.

• Proficiency in PowerShell scripting and automation preferred.

• Experience in a HIPAA-compliant setting preferred.

• Security experience beyond day-to-day operations, including red teaming or penetration testing, preferred.


🏝️ Benefits

• Flexible work schedule.

• Unlimited Paid Time Off (PTO).

• Wellness benefits for physical and mental health.

• Medical insurance coverage.

• Parental leave provisions.

• 401K retirement plan.

• Company-sponsored events.

• Employee referral program.

• Onsite gym facilities.

• Dog-friendly office environment.

• Snacks provided in the office.

• Commuter benefits available.

• Onsite massage services.

People also viewed

Harris Computer1 day ago

Platform & DevSecOps Delivery Architect

US flagAlabama, +20 more statesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TheWhiteam1 day ago

DevOps Engineer

ES flagSpain OnlyFreelanceDevOps & Site Reliability Engineer (SRE)
ApplyView job
CFactory-Creations1 day ago

Senior Site Reliability Engineer – Eastern Europe

BG flagBulgaria, +2 more countriesFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Capgemini1 day ago

Mainframe DevOps Migration Consultant

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$76.7k – $150.8k/year
ApplyView job
URBN (Urban Outfitters, Anthropologie Group, Free People & Nuuly)1 day ago

Senior DevOps Engineer

US flagPennsylvania OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Nitrado1 day ago

Site Reliability Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers