
Senior Software Engineer, Fleet Intelligence Agent Systems
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in California, +1 more state.
• Design and create Fleet Intelligence agent software that operates on Linux hosts, bare metal systems, and Kubernetes GPU nodes.
• Develop telemetry and health monitoring systems for NVIDIA GPUs, including DCGM/NVML, drivers, CUDA runtime, InfiniBand, containers, kernel/OS states, CPU, memory, disk, and networking.
• Create workflows for inventory, enrollment, node identity, local state, attestation, and backend exports.
• Manage local API, Prometheus metrics, file exports, and OTLP/HTTP export pathways.
• Enhance Kubernetes DaemonSet deployment, systemd service packaging, .deb/.rpm packaging, and container image processes.
• Contribute code, tests, documentation, release artifacts, and community-oriented engineering practices to the open-source Fleet Intelligence agent and collector software.
• Assist in out-of-band data collection using Redfish/BMC interfaces for inventory, GPU attestation, BMC metrics, and log management.
• Refine collector concurrency, rate limiting, retry mechanisms, credential management, partial-failure handling, and backend submission protocols.
• Collaborate with backend, infrastructure, SRE, security, and data center operations teams to ensure dependable GPU fleet observability.
• Over 5 years of professional software engineering experience.
• Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
• Extensive experience in Go development for Linux services, CLIs, agents, or daemons.
• Familiarity with Rust or a strong desire to work with Rust.
• In-depth knowledge of Linux systems, including processes, filesystems, networking, service lifecycle, logs, permissions, and host diagnostics.
• Experience in telemetry, observability, health monitoring, or fleet management systems.
• Proficient in Docker, Kubernetes, Helm, and production deployment processes.
• Experience contributing to open-source projects or engaging in public repositories with code reviews, issue tracking, documentation, release notes, and signed commits.
• Understanding of secure credential management, enrollment processes, tokens/JWTs, and service-to-service authentication.
• Strong debugging capabilities across hardware-adjacent software, operating systems, containers, and distributed backend integrations.
• Familiarity with NVIDIA datacenter GPUs, DGX systems, DCGM, NVML, CUDA, GPU drivers, XID/SXID events, or GPU diagnostics.
• Experience in developing host agents, node agents, collectors, monitoring daemons, or Kubernetes daemonsets.
• Background in Redfish, BMCs, firmware inventory, secure boot state, PCIe device inventory, or hardware attestation.
• Knowledge of OpenTelemetry, Prometheus, OTLP gateways, or metric/log export pipelines.
• Experience operating software in AI, HPC, cloud, or large-scale data center settings.
• Proven track record of significant contributions to open-source projects in systems software, observability, Kubernetes, Linux, hardware telemetry, Rust, or Go ecosystems.
• Equity
• Benefits
Bet On Talent
CVS Health
Scale Army Careers
WEX
Get handpicked remote jobs straight to your inbox weekly.