
Member of Technical Staff – Infrastructure
Posted Jul 25

Posted Jul 25
This is a fully remote position, open to applicants in United States.
• Ensure that deployments are smooth and uneventful (in the best way possible)
• Take ownership of CI/CD pipelines: enhance build times, improve caching strategies, and minimize flakiness
• Advance our Kubernetes (EKS) deployment approach for increased reliability and speed
• Develop and strengthen the infrastructure that supports model serving, inference, and agent tooling—not just the application layer
• Enhance our telemetry with improved instrumentation, intelligent sampling, and actionable dashboards, including evaluation pipelines and LLM-ops guardrails
• Create alerting mechanisms that identify real issues while filtering out irrelevant noise
• Accelerate the feedback loop from code to production
• Enhance preview environments, local development tools, and testing infrastructure
• Reduce manual effort through strategic automation, avoiding the creation of yet another dashboard that goes unnoticed
• Be the engineer who empowers other engineers—and agents—to work more efficiently
• Have experience from a company where infrastructure was the core product—not merely infrastructure support for another product.
• Applied your infrastructure expertise specifically to AI/LLM workloads: model serving, inference infrastructure, agent tooling, evaluation pipelines, or LLM-ops guardrails—working at an "AI company" is insufficient if the infrastructure work lacks this perspective
• Possess significant technical depth in distributed systems: Rust or Go, storage engines, control planes, Ceph, RDMA, eBPF, bare-metal automation, or Kubernetes internals (not just usage)
• A proven track record of exceptional accomplishments and impact
• Strong skills in Terraform—you have managed real infrastructure as code
• Hands-on experience with observability tools: OpenTelemetry, Datadog, Dash0, Braintrust, distributed tracing, metrics, structured logging
• Have been on-call and have developed systems that improve the on-call experience
• Think like a product manager for internal tools, where the focus is on enhancing developer (and agent) productivity
• Willingness to work diligently, move quickly, and adapt rapidly in a fast-paced environment
• A humble demeanor, a willingness to assist colleagues, and a commitment to doing whatever it takes to ensure team success
• Nice to have:
• A personal or self-directed infrastructure background: side projects, homelabs, open-source infrastructure tooling, published writing or talks—demonstrating genuine intrinsic interest, not just job experience
• Security expertise: IAM, zero-trust, secrets management
• Familiarity with SRE practices: SLOs, SLIs, error budgets, chaos engineering
• Experience in cost optimization for cloud infrastructure
• Based in Atlanta (preferred, but not mandatory)
• Offers Equity
Get handpicked remote jobs straight to your inbox weekly.