Remotery

Site Reliability Engineer

Posted Aug 4

This is a fully remote position, open to applicants in California.

📋 Description

• Provide support for the Veeam Data Cloud SaaS platform’s Government and Sovereign Cloud environment.

• Map out platform systems, workloads, dependencies, and areas of risk.

• Collaborate with subject-matter experts to address knowledge gaps and create onboarding materials.

• Draft and maintain runbooks, architectural documentation, and operational guides.

• Design highly available and fault-tolerant infrastructure on Azure, including Azure Government.

• Define Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.

• Lead incident responses and conduct blameless postmortems, transforming incidents into opportunities for improvement.

• Identify reliability risks and develop remediation strategies that adhere to compliance requirements.

• Establish requirements for observability instrumentation and lead its implementation.

• Set standards for alerting, telemetry, and monitoring.

• Create automation solutions to minimize manual work and support fleet management.

• Participate in on-call rotation duties.

• Work with infrastructure as code, CI/CD, deployment automation, and configuration management within air-gapped or compliance-constrained environments.

• Develop and maintain testing, canary deployment, and release validation pipelines.

• Integrate chaos engineering practices and monitoring tools.

• Collaborate with product, platform, security, legal, compliance, and operations teams.

• Take ownership of reliability issues from end to end and drive effective solutions.

• Mentor engineers and promote SRE practices throughout the organization.


⛳️ Requirements

• 7+ years of experience in Software Engineering, including a minimum of 3 years in Site Reliability Engineering (SRE), Platform Engineering, or related roles.

• Proven experience across multi-service platforms.

• Familiarity with Government or Sovereign Cloud environments, such as Azure Government or AWS GovCloud.

• Experience in regulated compliance environments, including but not limited to FedRAMP, CMMC, IL2/IL4/IL5, PCI-DSS, SOX, HIPAA, or HITRUST.

• Strong background in building and managing production services on cloud infrastructure, preferably Azure, including Azure Government.

• Capability to rapidly learn complex platforms with minimal guidance and limited access to restricted environments.

• Proficient in independently investigating systems and generating clear documentation, risk assessments, and improvement plans.

• Experience programming in TypeScript/JavaScript, Go, Java, C#, or equivalent languages.

• Familiarity with monitoring and observability tools such as Prometheus, Grafana, OpenTelemetry, or the ELK Stack.

• Experience with infrastructure as code tools, including Terraform, Terragrunt, or Pulumi.

• Knowledge of container orchestration, particularly with Kubernetes.

• Proficiency in CI/CD and GitOps tools, including GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or Dagger.

• Strong understanding of distributed systems, networking, and cloud-native architecture.

• Excellent written and verbal communication skills.


🏝️ Benefits

• Unlimited paid time off.

• 12 paid holidays, including 4 global VeeaMe Days dedicated to self-care.

• 24 paid volunteer hours each year through Veeam Cares.

• Paid parental leave: 8 weeks for all parents, 16 weeks for birthing parents.

• Comprehensive medical, dental, and vision coverage starting on the first day of employment.

• Mental health support, therapy sessions, and digital wellness resources available through the Employee Assistance Program.

• 401(k) retirement plan with matching contributions from the company.

• Support for fertility, adoption, and surrogacy through Maven.

• AirVet: 24/7 virtual veterinary care at no cost.

• Access to legal services, identity protection, and supplemental health insurance options.

• Tax-advantaged spending accounts for healthcare, dependent care, and commuting expenses.

• On-demand learning resources, mentoring opportunities, workshops, and learning events, including the annual Global Day of Learning.

• Competitive salary and benefits package.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers