
Staff MaaS Backend Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California, +1 more state.
• Collaborate with the Principal Architect on the comprehensive design of the MaaS system.
• Create decision records and advocate for technical choices.
• Lead the platform maturity roadmap by developing operable, measurable, and independently deployable capabilities.
• Establish Go and API standards, guide engineers, and coordinate with cross-functional teams.
• Oversee wire compatibility for OpenAI and Anthropic formats, which includes streaming, tool calling, and structured output.
• Advance routing functionality with load-aware, model-aware, and prefix-cache-aware endpoint selection, as well as circuit breaking and fallback strategies.
• Manage API versioning and deprecation agreements.
• Enhance token throughput, KV cache tiering, and optimize p95/p99 TTFT latency.
• Incorporate deployment tools, LoRA multiplexing, and cold-start-aware autoscaling with Kubernetes.
• Scale infrastructure to accommodate regional inference pools, capacity-aware failovers, and an active-active control plane.
• Define and uphold Service Level Objectives.
• Perform peak-concurrency load testing, maintain on-call documentation, and implement load shedding measures.
• Enforce policies for authorization, tenant isolation, zero-retention data paths, API key and OAuth credential lifecycle management, as well as distributed rate limits and quotas, and controls against abuse and prompt injection.
• Develop idempotent exactly-once token metering and a high-volume usage ledger with pre-paid spending caps and invoice reconciliation.
• Carry out incremental zero-regression migrations using strangler replacements, shadow traffic, and stateful dual-writes.
• Provide request tracing and cost telemetry through internal token and quality-attribute schemas.
• Ensure prompts, completions, and PII are not stored beyond explicitly consented policies.
• Take ownership of the implementation from start to finish, including code, migrations, SLOs, and on-call responsibilities for a revenue-generating service.
• Over 8 years of backend engineering experience.
• More than 3 years of experience managing a high-traffic, multi-tenant API platform for paying clients.
• Ability to work closely with architects and produce detailed design documents.
• Expertise in multi-region active-active distributed systems, caching, backpressure, and performance optimization.
• Proven ability to execute zero-downtime brownfield migrations for stateful subsystems without regressions or accounting discrepancies.
• Extensive hands-on experience with Go-based services.
• Production experience with Kubernetes, including Envoy and GPU-aware scheduling.
• Familiarity with the architectural trade-offs of PostgreSQL, Redis, and Kafka.
• Experience in developing end-to-end observability strategies using OpenTelemetry and high-cardinality analytics stores.
• Systems-level comprehension of LLM serving, including server-sent event streaming, KV/prefix caching, and TTFT/throughput trade-offs.
• Background in building reliable exactly-once metering and billing systems that reconcile billions of events during partial failures.
• Capacity to enforce fail-closed authorization, verified identities, and cross-tenant isolation.
• Operational maturity, which includes on-call responsibilities and conducting blameless incident reviews.
• Equal employment opportunities in compliance with country, state, and local regulations.
• Options for remote work within San Jose, CA, or Austin, TX locations.
Sigma Software Group
Plain Concepts
GitLab
Get handpicked remote jobs straight to your inbox weekly.