Remotery

Senior Site Reliability Engineer

Posted Aug 4

This is a fully remote position, open to applicants in California.

📋 Description

• Develop reliability practices for Veeam Data Cloud’s Government and Sovereign Cloud environment.

• Identify platform systems, dependencies, workloads, and risk factors.

• Produce onboarding materials, runbooks, architecture documentation, and operational manuals.

• Architect highly available and fault-tolerant infrastructure on Azure and Azure Government.

• Establish SLIs, SLOs, and error budgets.

• Oversee incident response and conduct blameless postmortems.

• Detect reliability risks and devise remediation strategies within compliance frameworks.

• Specify observability instrumentation needs and lead their implementation.

• Set up alerting, telemetry, and monitoring standards.

• Create automation to minimize toil and assist in fleet management.

• Engage in on-call rotations.

• Work with IaC, CI/CD, deployment automation, and configuration management in air-gapped or restricted settings.

• Construct and maintain testing, canary deployment, and release validation pipelines.

• Integrate chaos engineering and monitoring tools.

• Collaborate with product, platform, security, legal, compliance, and operations teams.

• Take ownership of reliability challenges from start to finish and drive solutions.

• Mentor engineers and promote SRE practices throughout Veeam.


⛳️ Requirements

• Over 7 years of experience in Software Engineering, with at least 3 years in SRE, Platform Engineering, or a similar role, across multi-service platforms.

• Familiarity with Government or Sovereign Cloud, such as Azure Government or AWS GovCloud.

• Experience in regulated compliance settings, including FedRAMP, CMMC, IL2/IL4/IL5, PCI-DSS, SOX, HIPAA, or HITRUST.

• Extensive experience building and managing production services on cloud infrastructure; Azure experience preferred, including Azure Government.

• Capacity to rapidly understand large, complex platforms with limited guidance and restricted access to environments.

• Proficiency in independently investigating systems and creating clear documentation, risk assessments, and improvement plans.

• Programming expertise in TypeScript/JS, Go, Java, C#, or similar languages.

• Experience with monitoring and observability tools such as Prometheus, Grafana, OpenTelemetry, or the ELK stack.

• Familiar with IaC tools like Terraform, Terragrunt, or Pulumi.

• Experience in container orchestration using Kubernetes.

• Proficiency in CI/CD and GitOps tools such as GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or Dagger.

• Strong grasp of distributed systems, networking, and cloud-native architecture.

• Excellent written and verbal communication skills.

• Bonus: Experience with B2B SaaS platforms in regulated or government markets.

• Bonus: Background in chaos engineering, resilience testing, or performance/load testing.

• Bonus: Experience in establishing an SRE or reliability function from the ground up.

• Bonus: Experience across modern cloud-native and legacy systems.

• Bonus: Familiarity with AI-first development workflows using LLM-powered tools.


🏝️ Benefits

• Unlimited paid time off.

• 12 paid holidays, including 4 global VeeaMe Days for self-care.

• 24 paid volunteer hours each year through Veeam Cares.

• Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents.

• Medical, dental, and vision coverage effective from the first day.

• Mental health support, therapy sessions, and digital wellness tools through the Employee Assistance Program.

• 401(k) retirement plan with company matching contributions.

• Fertility, adoption, and surrogacy support via Maven.

• AirVet 24/7 virtual veterinary care at no charge.

• Legal services, identity protection, and supplemental health insurance options.

• Tax-advantaged spending accounts for healthcare, dependent care, and commuting.

• On-demand learning resources, including LinkedIn Learning and O’Reilly.

• Mentoring, workshops, and educational events, including the annual Global Day of Learning.

• Competitive performance-based bonus included in total target compensation.

• Comprehensive benefits package encompassing health coverage, retirement plans, and unlimited time off.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers