
Technical Product Manager, Observability
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in United States.
• Take charge of the vision, roadmap, and priorities for k0rdent AI observability encompassing GPU compute, east-west fabric, high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services.
• Convert requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into a clear product direction.
• Collaborate with engineering to define requirements and assess trade-offs.
• Oversee the observability backlog by incorporating feedback from production deployments and design partners.
• Monitor and influence the response to evolving observability standards and technologies, including OpenTelemetry, DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling.
• Establish integration strategies for vendor telemetry sources into a cohesive, operator-facing observability platform.
• Work alongside product marketing and field teams on positioning, technical briefs, and reference architectures.
• Represent Mirantis to customers, analysts, and ecosystem partners.
• A minimum of 5 years in product management or a senior technical position overseeing an observability product or managing large-scale monitoring infrastructure.
• Proficient understanding of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch).
• Expertise in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture.
• Capability to collaborate directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals.
• Familiarity with GPU observability, including DCGM metrics, AI workload profiling, and performance analysis.
• Understanding of east-west fabric telemetry, such as InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health.
• Experience with high-performance storage telemetry from platforms like VAST Data, Weka, or DDN, encompassing IOPS, latency, and throughput instrumentation at scale.
• Knowledge of NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability.
• Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines.
• Employees have the option to work remotely.
• Collaborate with a top-notch, distributed team dedicated to openness and technical excellence.
• Establish the observability foundation for the AI cloud era, directly engaging with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises.
• Shape the product narrative and drive go-to-market success.
GE Vernova
DEUNA
Nagarro
Aspire Software
Get handpicked remote jobs straight to your inbox weekly.