
Senior Cloud Infrastructure, DevOps Solutions Architect
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in France, +3 more countries.
β’ Take ownership of comprehensive validation for partner software stacks, which includes cluster-wide stability testing, acceptance of real training workloads, and multi-day, multi-rack burn-in.
β’ Reduce the time taken from cluster handover to the initiation of the first production workload by effectively coordinating hardware bring-up, managed-service intake, and partner operations teams.
β’ Manage Day 2 production stability at scale across fleets, including aspects such as monitoring, logging, workload orchestration, fault detection and remediation, preventive maintenance, and the rollout of firmware and field-notice campaigns.
β’ Evaluate customer environments and operate diverse open platforms including Kubernetes, KubeVirt, Slurm, and GPU-aware schedulers.
β’ Integrate enterprise-grade networking and storage solutions while enabling third-party ISV workloads.
β’ Deliver consultative support and hands-on troubleshooting across various domains including bare metal, operating systems, software stacks, container platforms, networking, and storage.
β’ Assist R&D initiatives, proofs of concept, and proofs of value by validating new features, architectures, and upgrade strategies.
β’ Serve as the technical leader for designated accounts.
β’ Conduct structured knowledge transfer and enablement sessions.
β’ Create runbooks, onboarding documentation, and best-practice guides for partner teams.
β’ BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.
β’ Over 8 years of experience in managing scalable cloud environments and roles in automation engineering.
β’ Solid understanding of networking principles and data center architectures.
β’ Practical experience in managing HPC/AI clusters and NVIDIA GPU-accelerated infrastructure, encompassing deployment, driver and CUDA toolkit management, optimization, workload profiling, and troubleshooting.
β’ Extensive experience with Kubernetes for container orchestration, resource scheduling, and scaling in GPU-accelerated and HPC settings.
β’ Familiarity with scheduler internals, batch schedulers like Slurm, and mixed bare-metal/virtualized multi-tenant environments such as KubeVirt.
β’ In-depth knowledge of Linux, including RedHat and Ubuntu, OS-level security, and network protocols.
β’ Experience with storage solutions including Lustre, GPFS, ZFS, XFS, and Kubernetes storage technologies.
β’ Proficient in Python and Bash scripting.
β’ Knowledge of configuration management and Infrastructure-as-Code tools like Ansible and Terraform.
β’ Experience with GitOps-based management for cluster lifecycle and upgrades across large fleets.
β’ Familiarity with observability stacks such as Grafana, Loki, and Prometheus.
β’ Capability to assess and enhance MTBI and job goodput on extensive GPU clusters.
β’ Experience with fault detection, drain and remediation workflows, SLO/error-budget definitions, and post-incident reviews.
β’ Strong consultative background in leading architectural reviews and presenting to executive stakeholders.
β’ Knowledge of CI/CD pipelines and container-based microservices architectures.
β’ Experience with NVIDIA GPU and Network Operators, as well as NVIDIA Base Command Manager.
β’ Familiarity with DCGM, XID diagnostics, node-level health agents, and fleet-wide reliability intelligence.
β’ Expertise in AI-native scheduling and inference frameworks on Kubernetes, such as KAI, Grove, Dynamo, and NVIDIA Cloud Functions.
β’ Background with RDMA-based fabrics like InfiniBand or RoCE.
β’ Exposure to Cumulus Linux, SONiC, Spectrum-X fabrics, DPU/DOCA infrastructure services, and NVLink/NVSwitch partition operations is considered a strong advantage.
β’ Work with NVIDIA products and contribute to next-generation AI/HPC systems.
β’ Gain exposure to large-scale infrastructure projects and advanced GPU/HPC technologies.
β’ Collaborate with customers, partners, and cross-functional teams.
β’ Engage in technical leadership, knowledge transfer, and enablement opportunities.
Expel
Grafana Labs
Livestock Information Ltd
Salesforce
Get handpicked remote jobs straight to your inbox weekly.