Site Reliability Engineer, Provider Operations

Posted 23 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Develop and maintain monitoring systems for each provider and endpoint, focusing on latency, throughput, error rates, uptime, and output accuracy.

• Establish SLOs for each provider tier and generate actionable alerts.

• Enhance the identification of degraded endpoints and collaborate with the routing team for automatic traffic failover.

• Manage on-call responsibilities for provider incidents, which include triage, mitigation, communication with providers, conducting postmortems, and ensuring follow-up closure.

• Produce provider scorecards and SLO reports.

• Act as the technical escalation point for underperforming provider endpoints.

• Develop continuous canaries and evaluations to identify silent quality regressions.

• Automate tasks related to provider operations, such as disabling endpoints, adjusting capacity, managing deprecations, and tuning rate limits.

• Create endpoint load-testing tools to ensure launch readiness and handle day-zero traffic.

• Report directly to the Provider Operations Manager.


⛳️ Requirements

• Minimum of 4 years in SRE, production engineering, or infrastructure roles managing high-traffic, customer-facing systems.

• Strong expertise in observability practices, including metrics, tracing, logs, SLOs/error budgets, and actionable alerting.

• Proficient software engineer with a preference for developing tools rather than following runbooks.

• Experience with TypeScript and/or Python.

• Familiarity with failure modes in distributed systems, such as timeouts, retries, backpressure, and partial outages.

• Ability to act as a composed and clear incident commander while communicating with external partners under pressure.

• Understanding of, or willingness to learn about, LLM inference concepts including streaming, tool calling, prompt caching, throughput/latency trade-offs, and differences in provider APIs.

• Experience with an inference provider, model lab, GPU cloud, or API gateway/CDN company is a plus.

• Familiarity with TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, or Vercel is advantageous.

• Background in routing, load balancing, or traffic management systems is a plus.

• Experience with evaluations or synthetic monitoring for ML systems is a plus.


🏝️ Benefits

• [Detail any benefits offered such as health insurance, retirement plans, flexible work hours, etc.]

• [List any additional benefits like professional development opportunities, wellness programs, etc.]

People also viewed

Slate Auto19 hours ago

Senior Software Engineer, DevOps

US flagWashington OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$135.6k – $203.3k/year
ApplyView job
Funding Xchange22 hours ago

Senior DevOps/SRE Engineer

GB flagUnited Kingdom OnlyFull-timeDevOps & Site Reliability Engineer (SRE)£70k – £85k/year
ApplyView job
Leidos22 hours ago

DevOps Engineer – Technical Integration Lead

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$154.1k – $278.5k/year
ApplyView job
LeoLabs23 hours ago

Senior Site Reliability Engineer, SRE

US flagCalifornia OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$171k – $192k/year
ApplyView job
Mirantis1 day ago

Senior DevOps Engineer – PostgreSQL, Kafka, Kubernetes

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)
ApplyView job
LMI1 day ago

DevOps Engineer

US flagUnited States OnlyFull-timeDevOps & Site Reliability Engineer (SRE)$122k – $211k/year
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers