
Site Reliability Engineer, Provider Operations
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in United States.
• Develop and maintain monitoring systems for each provider and endpoint, focusing on latency, throughput, error rates, uptime, and output accuracy.
• Establish SLOs for each provider tier and generate actionable alerts.
• Enhance the identification of degraded endpoints and collaborate with the routing team for automatic traffic failover.
• Manage on-call responsibilities for provider incidents, which include triage, mitigation, communication with providers, conducting postmortems, and ensuring follow-up closure.
• Produce provider scorecards and SLO reports.
• Act as the technical escalation point for underperforming provider endpoints.
• Develop continuous canaries and evaluations to identify silent quality regressions.
• Automate tasks related to provider operations, such as disabling endpoints, adjusting capacity, managing deprecations, and tuning rate limits.
• Create endpoint load-testing tools to ensure launch readiness and handle day-zero traffic.
• Report directly to the Provider Operations Manager.
• Minimum of 4 years in SRE, production engineering, or infrastructure roles managing high-traffic, customer-facing systems.
• Strong expertise in observability practices, including metrics, tracing, logs, SLOs/error budgets, and actionable alerting.
• Proficient software engineer with a preference for developing tools rather than following runbooks.
• Experience with TypeScript and/or Python.
• Familiarity with failure modes in distributed systems, such as timeouts, retries, backpressure, and partial outages.
• Ability to act as a composed and clear incident commander while communicating with external partners under pressure.
• Understanding of, or willingness to learn about, LLM inference concepts including streaming, tool calling, prompt caching, throughput/latency trade-offs, and differences in provider APIs.
• Experience with an inference provider, model lab, GPU cloud, or API gateway/CDN company is a plus.
• Familiarity with TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, or Vercel is advantageous.
• Background in routing, load balancing, or traffic management systems is a plus.
• Experience with evaluations or synthetic monitoring for ML systems is a plus.
• [Detail any benefits offered such as health insurance, retirement plans, flexible work hours, etc.]
• [List any additional benefits like professional development opportunities, wellness programs, etc.]
Slate Auto
Funding Xchange
Leidos
LeoLabs
Get handpicked remote jobs straight to your inbox weekly.