Senior Site Reliability Engineer

Posted 15 hours ago

This is a fully remote position, open to applicants in New York.

📋 Description

• Define and manage the architectural framework of Chalice's AI platform.

• Report directly to the VP of Engineering and collaborate cross-functionally with teams in Engineering, Data Science, Machine Learning, and Product.

• Mentor and enhance the skills of engineers in infrastructure best practices, operational excellence, and architectural strategies.

• Design and manage scalable, event-driven, multi-tenant ML infrastructure.

• Facilitate distributed ML training utilizing Databricks, Ray, and Flyte on EKS.

• Deliver containerized solutions to external clients.

• Construct and maintain AWS event-driven systems using EventBridge, MSK/Kafka, Kinesis, Lambda/Fargate, SQS/SNS, and Step Functions.

• Design centralized state storage solutions using DynamoDB, Redis, and Postgres.

• Establish standards for idempotency, replay safety, event schema governance, and traceability.

• Operate multi-cluster Kubernetes environments in a production setting.

• Implement GitOps, progressive delivery, cluster-level security policies, and multi-tenant isolation.

• Develop standards for Terraform modules and reusable infrastructure components.

• Enforce GitHub best practices and enhance CI/CD pipelines with infrastructure testing, policy validation, progressive deployment, and rollback capabilities.

• Standardize federated identity management, Azure SSO, API authentication, key management, and resource isolation.

• Define external model-serving architecture, encompassing OCI packaging, secret injection, network isolation, telemetry, upgrades, and runtime contracts.

• Establish SLIs, SLOs, error budgets, distributed tracing, golden signals, on-call structures, escalation paths, incident response, and postmortem processes.

• Own the Databricks infrastructure, including SSO, authentication, provisioning, Terraform deployments, cluster policies, and Unity Catalog governance.

• Contribute to building and scaling the SRE function while establishing organizational reliability standards.


⛳️ Requirements

• 8–10 years of experience in designing and managing production infrastructure at scale.

• Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

• Extensive expertise in AWS architecture, covering networking, SSO, compute, storage, and event-driven services.

• Experience in managing Kubernetes control planes within production environments.

• Proven experience in building or migrating event-driven systems at scale.

• Background in ML-focused or data-intensive settings.

• Experience in replacing or removing legacy infrastructure components and streamlining system design.

• Experience in taking ownership of production failures and leading postmortems that drive systemic improvements.

• Strong architectural judgment with the ability to constructively challenge assumptions.

• Certifications such as AWS, Databricks, or Kubernetes are advantageous but not mandatory.


🏝️ Benefits

• Equity.

• Medical, dental, and vision insurance.

• 401(k) options.

• Unlimited PTO.

• 11 company holidays.

• Office-wide closure from Christmas Eve to New Year's.

• Office start-up stipend.

• In-office meal allowance.

• Opportunities for professional development and career advancement.

• An innovative, collaborative, and inclusive workplace.

People also viewed

DaVita Kidney Care15 hours ago

Senior DevOps Engineer

US flagTennessee OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$115k – $183k/year
ApplyView job
WEX16 hours ago

SRE & Application Services Intern – Graduate/Master's

US flagUnited States OnlyInternshipDevOps & Site Reliability Engineer (SRE)$30 – $45/hour
ApplyView job
WEX16 hours ago

DevOps Engineer, AI Engineering Intern

US flagUnited States OnlyInternshipDevOps & Site Reliability Engineer (SRE)
ApplyView job
Dscout16 hours ago

Senior DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
DaVita Kidney Care18 hours ago

Senior Manager, DevOps

US flagFlorida OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$115k – $183k/year
ApplyView job
Cint19 hours ago

Senior Cloud Infrastructure Engineer – DevOps

ES flagSpain OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers