
Senior Solutions Architect, Generative AI
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in California, +1 more state.
• Work closely with clients to optimize GPU usage and enhance end-to-end workload efficiency while boosting infrastructure reliability and minimizing costs.
• Design and refine large-scale AI clusters, focusing on GPU computing, high-performance networking, storage solutions, workload scheduling, orchestration, and observability.
• Analyze distributed training and inference tasks to pinpoint bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stacks.
• Troubleshoot intricate infrastructure and distributed systems challenges involving InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.
• Lead proof-of-concept initiatives and performance evaluations for extensive AI infrastructure, creating benchmarking tools, automation scripts, runbooks, and necessary technical documentation.
• Collaborate with NVIDIA’s engineering, product, and sales teams to achieve design wins and foster innovative solutions that align with customer needs and feedback from the field.
• Bachelor’s, Master’s, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or a related engineering discipline, or equivalent experience.
• Over 6 years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a comparable technical role.
• Profound knowledge of Linux systems, distributed computing, GPU architectures, and both hardware and software components of large-scale AI clusters.
• Practical experience in designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments utilizing technologies such as InfiniBand, RoCE, or GPUDirect RDMA.
• Background in debugging NCCL communication and distributed collective performance issues, including topology, transport, congestion, routing, and host-level configuration challenges.
• Skill in profiling AI workloads and detecting performance bottlenecks across computing, networking, storage, and orchestration layers.
• Familiarity with cluster schedulers and orchestration platforms like Kubernetes and Slurm, as well as experience with containers and production monitoring systems.
• Proficient in Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.
• Equity
• Generous benefits package
Snowflake
Cisco
BCD Travel
Get handpicked remote jobs straight to your inbox weekly.