
Senior Site Reliability Engineer
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in New York.
• Define and manage the architectural framework of Chalice's AI platform.
• Report directly to the VP of Engineering and collaborate cross-functionally with teams in Engineering, Data Science, Machine Learning, and Product.
• Mentor and enhance the skills of engineers in infrastructure best practices, operational excellence, and architectural strategies.
• Design and manage scalable, event-driven, multi-tenant ML infrastructure.
• Facilitate distributed ML training utilizing Databricks, Ray, and Flyte on EKS.
• Deliver containerized solutions to external clients.
• Construct and maintain AWS event-driven systems using EventBridge, MSK/Kafka, Kinesis, Lambda/Fargate, SQS/SNS, and Step Functions.
• Design centralized state storage solutions using DynamoDB, Redis, and Postgres.
• Establish standards for idempotency, replay safety, event schema governance, and traceability.
• Operate multi-cluster Kubernetes environments in a production setting.
• Implement GitOps, progressive delivery, cluster-level security policies, and multi-tenant isolation.
• Develop standards for Terraform modules and reusable infrastructure components.
• Enforce GitHub best practices and enhance CI/CD pipelines with infrastructure testing, policy validation, progressive deployment, and rollback capabilities.
• Standardize federated identity management, Azure SSO, API authentication, key management, and resource isolation.
• Define external model-serving architecture, encompassing OCI packaging, secret injection, network isolation, telemetry, upgrades, and runtime contracts.
• Establish SLIs, SLOs, error budgets, distributed tracing, golden signals, on-call structures, escalation paths, incident response, and postmortem processes.
• Own the Databricks infrastructure, including SSO, authentication, provisioning, Terraform deployments, cluster policies, and Unity Catalog governance.
• Contribute to building and scaling the SRE function while establishing organizational reliability standards.
• 8–10 years of experience in designing and managing production infrastructure at scale.
• Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
• Extensive expertise in AWS architecture, covering networking, SSO, compute, storage, and event-driven services.
• Experience in managing Kubernetes control planes within production environments.
• Proven experience in building or migrating event-driven systems at scale.
• Background in ML-focused or data-intensive settings.
• Experience in replacing or removing legacy infrastructure components and streamlining system design.
• Experience in taking ownership of production failures and leading postmortems that drive systemic improvements.
• Strong architectural judgment with the ability to constructively challenge assumptions.
• Certifications such as AWS, Databricks, or Kubernetes are advantageous but not mandatory.
• Equity.
• Medical, dental, and vision insurance.
• 401(k) options.
• Unlimited PTO.
• 11 company holidays.
• Office-wide closure from Christmas Eve to New Year's.
• Office start-up stipend.
• In-office meal allowance.
• Opportunities for professional development and career advancement.
• An innovative, collaborative, and inclusive workplace.
DaVita Kidney Care
WEX
WEX
Dscout
Get handpicked remote jobs straight to your inbox weekly.