
Staff AI Observability, Telemetry Engineer
Posted 7 hours ago

Posted 7 hours ago
This is a fully remote position, open to applicants in California, +1 more state.
• Design and expand a high-cardinality telemetry infrastructure utilizing robust time-series databases such as VictoriaMetrics, Thanos, or Mimir.
• Integrate hardware-level exporters, including NVIDIA DCGM, network switch telemetry, and IPMI/Redfish, into the Kubernetes observability framework.
• Create eBPF-based diagnostic tools to monitor network congestion, kernel-level I/O latency, and identify distributed training bottlenecks.
• Develop automated dashboards and alerting systems that proactively isolate degraded hardware before it affects customer training operations.
• Design metric pipelines for precise, multi-tenant billing based on real-time GPU and network utilization data.
• Collaborate with GPU Systems and Scheduling teams to establish observability standards for AI-native workloads.
• Lead technical design reviews pertaining to observability architecture.
• Mentor team members on best practices for high-performance telemetry collection and analysis.
• Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related discipline.
• Over 6 years of experience in software or site reliability engineering.
• Extensive, hands-on knowledge of the Prometheus/OpenTelemetry ecosystem.
• Advanced proficiency in the Go programming language.
• Significant experience in writing custom Kubernetes metric exporters and operators.
• Practical experience with kernel-level tracing tools, including eBPF and BCC.
• In-depth performance tuning experience with Linux systems.
• Strong understanding of AI hardware metrics, including GPU power states, SM utilization, and memory bandwidth.
• Proficient in high-performance network telemetry.
• Demonstrated history of operating, debugging, and scaling large-scale telemetry stacks in high-performance computing or cloud environments.
• Strong technical leadership abilities and capacity to influence architectural decisions while aligning cross-functional teams with observability standards.
• Exceptional communication skills and the ability to translate complex system requirements into achievable engineering milestones.
• Experience in fast-paced, high-growth engineering settings is highly preferred.
• Equal employment opportunities in accordance with country, state, and local laws.
• Non-discrimination protections based on race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.
Sambatech
Connection
TSMG Holding
Twilio
Get handpicked remote jobs straight to your inbox weekly.