
Senior Software Engineer β AI Inference & Runtime Platform
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in Washington.
β’ Take ownership of the inference control plane for open-weight models and custom task-model zoos across managed GPU clouds and customer-managed Kubernetes clusters.
β’ Oversee the deployment and configuration of serving-tier engines, implement cold-start strategies, define per-model SLOs, and manage upgrade/canary processes.
β’ Administer the Kubernetes infrastructure for inference and sandbox workloads, which includes operators, CRDs, autoscaling, GPU scheduling and sharing, and node lifecycle management.
β’ Manage the stateful data plane for hosted vector stores and graph stores, including deployment, backups, scaling, recovery, and verified restores.
β’ Supervise the sandbox runtime and host-side control plane, focusing on lifecycle management, execution, snapshot/fork processes, teardown, metering, and isolation-boundary threat modeling.
β’ Lead FastAPI control-plane services, Terraform/OpenTofu, Bicep, and dashboards.
β’ Operate the platform layer that supports backend services and troubleshoot backend services as necessary.
β’ Guide the open-source initiative through the review of PRs, documentation, and ensuring reproducible builds.
β’ Develop the fork engine, guest agent, and multi-substrate model lifecycle utilizing Rust, Kubernetes controllers, and FastAPI endpoints.
β’ Precisely account for per-token costs, including scenarios involving client mid-stream disconnections.
β’ Facilitate the operation of hardware-isolated microVMs for secure and compliant execution of untrusted agent-generated code.
β’ A minimum of 5 years of experience in delivering production systems using a systems programming language.
β’ Proficiency in Rust is essential; strong experience in Go, C/C++, or Zig with a genuine interest in Rust is acceptable.
β’ Understanding of asynchronous runtimes, memory-safety practices, and debugging at the syscall boundary.
β’ Proven experience managing Kubernetes workloads, including controllers or operators, scheduling, autoscaling, and node lifecycle management.
β’ Capability to threat-model isolation boundaries, which includes namespaces, cgroups, seccomp, and hypervisors.
β’ Experience in implementing security best practices for agentic execution: ensuring least privilege, avoiding credentials in the sandbox, maintaining audit trails, and requiring human approval for write actions.
β’ Practical experience in deploying or managing open-weight LLM serving infrastructure, such as vLLM/SGLang or comparable solutions.
β’ Strong performance discipline in distributed systems.
β’ In-depth knowledge of Rust (tokio), Python/FastAPI, Kubernetes operators, KEDA, Karpenter, and GPU device plugins/DRA.
β’ Familiarity with isolation technologies like Firecracker, Kata, gVisor, or similar.
β’ Knowledge of secrets management and egress control, including tools like Vault/KMS.
β’ Experience in hosting stateful systems such as vector stores and graph stores, with a focus on backup and failover protocols.
β’ Comfort with operating across AWS, Azure, GCP, and managed GPU clouds.
β’ Familiarity with Terraform/OpenTofu, Bicep, and OpenTelemetry.
β’ A Bachelor's Degree is required; a Master's Degree is a plus.
β’ Must currently be authorized to work in the United States on a full-time basis.
β’ Willingness to travel twice a year for company summits is required.
β’ Competitive compensation package typical of early-stage startups, determined by capabilities, experience, and location.
β’ Eligibility for bonuses.
β’ Health insurance with significant coverage for dependents.
β’ Flexible paid time off policy.
β’ Equity options available.
β’ A fully remote work culture with a team based in Seattle.
β’ Company summits held twice a year.
Intermedia Cloud Communications
Juniper Square
EVERSANA
Get handpicked remote jobs straight to your inbox weekly.