Remotery

Senior Site Reliability Engineer

Posted 5 days ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the Databricks and Snowflake platform lifecycle, which encompasses automation, workspace governance, job orchestration, and cost optimization.

• Design and implement resilient, scalable, and secure infrastructure across various cloud environments.

• Lead initiatives related to failover, autoscaling, chaos testing, and capacity planning.

• Develop and sustain monitoring, alerting, and logging infrastructure utilizing Datadog and other open-source tools.

• Establish and enforce service level objectives (SLOs) and service level agreements (SLAs) for essential services.

• Automate the deployment of data pipelines, ML workflows, and infrastructure components using GitHub Actions, Terraform, and other Infrastructure as Code (IaC) tools.

• Create frameworks and tools for data movement across Snowflake, S3, Delta Lake, and Kafka, both inter- and intra-cloud.

• Leverage EventBridge, SNS/SQS, and Lambda to construct loosely coupled, scalable data systems.

• Collaborate with analytics, data science, product, and engineering teams.

• Shape decisions regarding data platform architecture, machine learning enablement, and data product strategies.


⛳️ Requirements

• A minimum of 6 years of experience in Site Reliability Engineering (SRE), platform engineering, or DevOps, particularly focusing on data-intensive or machine learning-driven applications.

• Daily proficiency with AI coding tools such as Claude Code, Cursor, Copilot, or similar.

• Hands-on experience with Databricks, specifically in workspace configuration, cluster/job management, and CI/CD and data orchestration integration.

• Familiarity with Snowflake.

• In-depth knowledge of AWS or comparable cloud-native infrastructure, including VPCs, IAM, event-driven patterns, and serverless computing.

• Proficiency in observability tools, particularly Datadog.

• Strong expertise in CI/CD tools, especially GitHub Actions.

• Experience with Infrastructure as Code, particularly Terraform.

• Basic knowledge of shell scripting and Python.

• Proven experience in building and maintaining highly available, fault-tolerant systems.

• Exceptional communication and collaboration abilities.

• No employment sponsorship is available.


🏝️ Benefits

• Comprehensive total rewards strategy.

• Post-offer health screenings and vaccinations as mandated by clients.

• Reasonable accommodations for individuals with physical and mental disabilities.

• Equal Employment Opportunity protections.

People also viewed

DATAGROUP2 days ago

DevOps Engineer

DE flagGermany OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
Ambush2 days ago

DevOps Engineer

BR flagBrazil OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DuoKey2 days ago

DevOps Engineer

MU flagMauritius OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
TEKsystems3 days ago

SRE – CloudOps, Practice Architect II

US flagIllinois OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
TEKsystems3 days ago

SRE CloudOps Practice Architect II

US flagTexas OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$148.2k – $222.4k/year
ApplyView job
Level Data3 days ago

Senior DevOps Engineer

US flagMassachusetts OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$120k – $135k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers