Remotery

Staff AI Observability, Telemetry Engineer

Posted 7 hours ago

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Design and expand a high-cardinality telemetry infrastructure utilizing robust time-series databases such as VictoriaMetrics, Thanos, or Mimir.

• Integrate hardware-level exporters, including NVIDIA DCGM, network switch telemetry, and IPMI/Redfish, into the Kubernetes observability framework.

• Create eBPF-based diagnostic tools to monitor network congestion, kernel-level I/O latency, and identify distributed training bottlenecks.

• Develop automated dashboards and alerting systems that proactively isolate degraded hardware before it affects customer training operations.

• Design metric pipelines for precise, multi-tenant billing based on real-time GPU and network utilization data.

• Collaborate with GPU Systems and Scheduling teams to establish observability standards for AI-native workloads.

• Lead technical design reviews pertaining to observability architecture.

• Mentor team members on best practices for high-performance telemetry collection and analysis.


⛳️ Requirements

• Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related discipline.

• Over 6 years of experience in software or site reliability engineering.

• Extensive, hands-on knowledge of the Prometheus/OpenTelemetry ecosystem.

• Advanced proficiency in the Go programming language.

• Significant experience in writing custom Kubernetes metric exporters and operators.

• Practical experience with kernel-level tracing tools, including eBPF and BCC.

• In-depth performance tuning experience with Linux systems.

• Strong understanding of AI hardware metrics, including GPU power states, SM utilization, and memory bandwidth.

• Proficient in high-performance network telemetry.

• Demonstrated history of operating, debugging, and scaling large-scale telemetry stacks in high-performance computing or cloud environments.

• Strong technical leadership abilities and capacity to influence architectural decisions while aligning cross-functional teams with observability standards.

• Exceptional communication skills and the ability to translate complex system requirements into achievable engineering milestones.

• Experience in fast-paced, high-growth engineering settings is highly preferred.


🏝️ Benefits

• Equal employment opportunities in accordance with country, state, and local laws.

• Non-discrimination protections based on race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

People also viewed

Sambatech5 hours ago

AI Specialist Fellow — LLMs & RAG, Product & Applied Research

BR flagBrazil OnlyFreelanceArtificial IntelligenceR$5,720/month
ApplyView job
Connection6 hours ago

AI Power Platform Administrator

US flagNew Hampshire OnlyFull-timeArtificial Intelligence$122.5k – $159.2k/year
ApplyView job
TSMG Holding6 hours ago

Quality Control Specialist, AI/ML – Greek

GR flagGreece OnlyFull-timeArtificial Intelligence
ApplyView job
Twilio6 hours ago

Senior Data and AI Specialist

US flagUnited States OnlyFull-timeArtificial Intelligence$106.3k – $156.3k/year
ApplyView job
Twilio6 hours ago

Senior Data and AI Specialist

CA flagCanada OnlyFull-timeArtificial IntelligenceC$99.8k – C$124.7k/year
ApplyView job
Bitdeer Group7 hours ago

Staff AI Scheduling, Orchestration Engineer

US flagCalifornia, +1 more stateFull-timeArtificial Intelligence
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers