Remotery

Senior Solutions Architect, Generative AI

Posted Jul 31

This is a fully remote position, open to applicants in California, +1 more state.

📋 Description

• Work closely with clients to optimize GPU usage and enhance end-to-end workload efficiency while boosting infrastructure reliability and minimizing costs.

• Design and refine large-scale AI clusters, focusing on GPU computing, high-performance networking, storage solutions, workload scheduling, orchestration, and observability.

• Analyze distributed training and inference tasks to pinpoint bottlenecks across GPUs, CPUs, memory, network fabrics, storage systems, and software stacks.

• Troubleshoot intricate infrastructure and distributed systems challenges involving InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch.

• Lead proof-of-concept initiatives and performance evaluations for extensive AI infrastructure, creating benchmarking tools, automation scripts, runbooks, and necessary technical documentation.

• Collaborate with NVIDIA’s engineering, product, and sales teams to achieve design wins and foster innovative solutions that align with customer needs and feedback from the field.


⛳️ Requirements

• Bachelor’s, Master’s, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or a related engineering discipline, or equivalent experience.

• Over 6 years of experience in AI infrastructure, systems engineering, high-performance computing, networking, site reliability engineering, or a comparable technical role.

• Profound knowledge of Linux systems, distributed computing, GPU architectures, and both hardware and software components of large-scale AI clusters.

• Practical experience in designing, deploying, operating, or troubleshooting high-performance GPU networks in on-premises or cloud environments utilizing technologies such as InfiniBand, RoCE, or GPUDirect RDMA.

• Background in debugging NCCL communication and distributed collective performance issues, including topology, transport, congestion, routing, and host-level configuration challenges.

• Skill in profiling AI workloads and detecting performance bottlenecks across computing, networking, storage, and orchestration layers.

• Familiarity with cluster schedulers and orchestration platforms like Kubernetes and Slurm, as well as experience with containers and production monitoring systems.

• Proficient in Python, shell scripting, or similar languages for infrastructure automation, benchmarking, and systems troubleshooting.


🏝️ Benefits

• Equity

• Generous benefits package

People also viewed

Whippy20 hours ago

Solutions Engineer

US flagUnited States OnlyFull-timeSolutions Engineer
ApplyView job
Snowflake22 hours ago

Senior Solution Engineer

US flagFlorida, +3 more statesFull-timeSolutions Engineer$138k – $181.1k/year
ApplyView job
Cisco22 hours ago

Partner Solutions Engineer

US flagNew Jersey OnlyFull-timeSolutions Engineer$212.2k – $268.1k/year
ApplyView job
BCD Travel1 day ago

AI Solutions Engineer – Architect

GB flagUnited Kingdom, +2 more statesFull-timeSolutions Engineer€110k – €190k/year
ApplyView job
AutoStore™1 day ago

Solutions Consultant – Warehouse Automation

AU flagAustralia OnlyFull-timeSolutions Engineer
ApplyView job
Agility Technologies Inc1 day ago

Solutions Architect, Databricks

US flagUnited States OnlyFull-timeSolutions Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers