
Technical Product Manager, Observability
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
• Take ownership of the vision, roadmap, and priorities for k0rdent AI observability across GPU compute, east-west fabric, high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services.
• Convert requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into a distinct product direction.
• Collaborate with engineering to establish requirements and assess trade-offs.
• Oversee the observability backlog using input from production deployments and design partners.
• Monitor and influence responses to emerging observability standards and technologies, including OpenTelemetry, DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling.
• Develop integration strategies for vendor telemetry sources into a cohesive, operator-facing observability plane.
• Work alongside product marketing and field teams on positioning, technical briefs, and reference architectures.
• Represent Mirantis in interactions with customers, analysts, and ecosystem partners.
• Over 5 years in product management or a senior technical position responsible for an observability product or managing large-scale monitoring infrastructure.
• Proficient understanding of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch).
• Strong knowledge of Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture.
• Capability to engage directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals.
• Experience with GPU observability, including DCGM metrics, AI workload profiling, and performance analysis.
• Familiarity with east-west fabric telemetry, such as InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health.
• Background in high-performance storage telemetry from platforms like VAST Data, Weka, or DDN, encompassing IOPS, latency, and throughput instrumentation at scale.
• Knowledge of NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability.
• Experience with workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines.
• Construct the observability foundation for the AI cloud era, collaborating directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises.
• Work alongside a world-class, distributed team dedicated to openness and technical excellence.
• Influence the product narrative and contribute to go-to-market success.
• Enjoy a remote work arrangement.
Abbott
UENI Ltd
Cencora
Appspace
Get handpicked remote jobs straight to your inbox weekly.