
Senior Site Reliability Engineer
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the Databricks and Snowflake platform lifecycle, which encompasses automation, workspace governance, job orchestration, and cost optimization.
• Design and implement resilient, scalable, and secure infrastructure across various cloud environments.
• Lead initiatives related to failover, autoscaling, chaos testing, and capacity planning.
• Develop and sustain monitoring, alerting, and logging infrastructure utilizing Datadog and other open-source tools.
• Establish and enforce service level objectives (SLOs) and service level agreements (SLAs) for essential services.
• Automate the deployment of data pipelines, ML workflows, and infrastructure components using GitHub Actions, Terraform, and other Infrastructure as Code (IaC) tools.
• Create frameworks and tools for data movement across Snowflake, S3, Delta Lake, and Kafka, both inter- and intra-cloud.
• Leverage EventBridge, SNS/SQS, and Lambda to construct loosely coupled, scalable data systems.
• Collaborate with analytics, data science, product, and engineering teams.
• Shape decisions regarding data platform architecture, machine learning enablement, and data product strategies.
• A minimum of 6 years of experience in Site Reliability Engineering (SRE), platform engineering, or DevOps, particularly focusing on data-intensive or machine learning-driven applications.
• Daily proficiency with AI coding tools such as Claude Code, Cursor, Copilot, or similar.
• Hands-on experience with Databricks, specifically in workspace configuration, cluster/job management, and CI/CD and data orchestration integration.
• Familiarity with Snowflake.
• In-depth knowledge of AWS or comparable cloud-native infrastructure, including VPCs, IAM, event-driven patterns, and serverless computing.
• Proficiency in observability tools, particularly Datadog.
• Strong expertise in CI/CD tools, especially GitHub Actions.
• Experience with Infrastructure as Code, particularly Terraform.
• Basic knowledge of shell scripting and Python.
• Proven experience in building and maintaining highly available, fault-tolerant systems.
• Exceptional communication and collaboration abilities.
• No employment sponsorship is available.
• Comprehensive total rewards strategy.
• Post-offer health screenings and vaccinations as mandated by clients.
• Reasonable accommodations for individuals with physical and mental disabilities.
• Equal Employment Opportunity protections.
DATAGROUP
Ambush
DuoKey
TEKsystems
Get handpicked remote jobs straight to your inbox weekly.