Remotery

Senior Site Reliability Engineer

Posted Aug 4

This is a fully remote position, open to applicants in United States.

📋 Description

• Act as the senior escalation point for intricate P1/P2 production incidents, managing cross-system triage and implementing permanent architectural solutions.

• Facilitate platform-level architecture assessments to ensure reliability, scalability, security, and operational compliance.

• Recognize systemic failure patterns and convert them into architectural modifications, design standards, and platform enhancements.

• Take ownership of the availability, reliability, performance, and scalability of production systems.

• Establish, monitor, and enhance SLOs, SLIs, and operational KPIs.

• Create and manage Terraform infrastructure-as-code solutions, including modules, state management, and governance frameworks.

• Reduce operational toil through automation, self-service functionalities, and platform tooling.

• Develop and sustain automation frameworks utilizing Bash, PowerShell, and other relevant scripting technologies.

• Oversee and design Microsoft Azure solutions, with AWS serving as a secondary platform.

• Manage Kubernetes in production settings, including cluster management and platform upkeep.

• Supervise hybrid-cloud environments, virtual machines, networking, and distributed infrastructure.

• Ensure Datadog observability and PagerDuty alerting configurations are maintained.

• Design infrastructure controls to achieve PCI-DSS, SOC 1/2, and ISO 27001 compliance and assist during audits.

• Collaborate with Engineering, Product, Security, and Operations teams on CI/CD, releases, and DevOps maturity initiatives.

• Guide engineers and impact organizational infrastructure design standards.


⛳️ Requirements

• Over 8 years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering.

• Direct responsibility for complex production platforms at scale.

• Senior technical escalation experience for cross-team incidents and architectural remediation efforts.

• In-depth Microsoft Azure infrastructure expertise with a working knowledge of AWS.

• Extensive experience with Kubernetes in production environments.

• Advanced skills in Terraform and Infrastructure-as-Code, including module creation, state management, and governance.

• Proficient in Bash scripting and operational automation development.

• Understanding of networking fundamentals including firewalls, DNS, routing, VPN, troubleshooting, and Cloudflare edge services.

• Experience with Active Directory administration and hybrid identity solutions.

• Background in CI/CD pipeline design and improvement of deployment workflows.

• Familiarity with Datadog, PagerDuty, or similar observability and alerting tools.

• Experience in compliance-regulated environments, such as PCI-DSS, SOC 1/2, and ISO 27001.

• Ability to influence across organizational boundaries without direct authority.

• Candidates must reside within the Eastern or Central time zones.

• Authorization to work in the U.S. without employer sponsorship is mandatory.

• Preferred: Experience with Puppet or similar configuration management tools.

• Preferred: Administration of Microsoft SQL Server.

• Preferred: Management of Meraki firewall policies.

• Preferred: Development of AI-driven operational workflows and Model Context Protocol (MCP).

• Preferred: Leadership in internal developer platforms or platform engineering initiatives.

• Preferred: Support for large-scale SaaS or high-availability platforms.

• Preferred: Knowledge of Scala and/or Java at the infrastructure level.

• Preferred: Certifications such as Azure Solutions Architect Expert, Azure Administrator Associate, AWS Solutions Architect, or CKA.


🏝️ Benefits

• 22 days of paid time off.

• 11 company-paid holidays.

• Medical insurance coverage.

• Dental insurance coverage.

• Vision insurance coverage.

• 5% 401(k) company match.

• A variety of other selectable benefits.

• Competitive salary.

• Opportunities for development and career advancement.

• Home-based/remote work arrangement available.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers