
Manager, Solutions Architecture – Emerging AI Labs
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in California, +2 more states.
• Recruit and oversee a team of solutions architects, system/network engineers, and software engineers dedicated to large-scale GPU and AI networking implementations.
• Establish priorities, allocate resources, mentor team members, and guarantee high-quality customer delivery across multiple concurrent projects.
• Engage directly in technical reviews, design decisions, and critical debugging initiatives.
• Offer subject-matter expertise in GPU and network systems, acting as the senior technical contact for strategic clients and emerging AI laboratories.
• Lead the configuration and performance debugging of compute and network systems to ensure reliable clusters.
• Facilitate discussions on compute, network, and software architecture while supporting server, network, and cluster initialization, which may include on-site data center tasks as necessary.
• Gather and synthesize requirements specific to customers.
• Collaborate with GPU and Network Systems Engineering, Product Management, and Sales to shape roadmaps and reference designs.
• Broaden NVIDIA solution engagements across key AI-lab clients.
• BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, or another Engineering discipline, or equivalent experience.
• Over 8 years of experience in Systems, Solutions, Field Engineering, Network Engineering, Data Center Engineering, or related roles.
• At least 2 years of experience in a leadership or mentoring capacity for engineers or architects.
• Expertise at the system level across CPU/GPU server architecture, NICs, Linux, system software, and kernel drivers.
• Proficient with Ethernet and/or InfiniBand switches, NICs, fabrics, relevant tooling, and cluster performance troubleshooting.
• Knowledgeable about data center infrastructure, including power, cooling, and deployment limitations.
• Capability to lead technical teams, establish priorities, and manage complex projects from design to production.
• Experience collaborating with Product Management, Sales, and Engineering teams.
• Strong time management skills with the ability to balance planning and hands-on support.
• Exceptional written and verbal communication abilities, including customer meetings, status and risk updates, design documents, debugging summaries, and presentations.
• Direct experience in people management and recruitment for geographically distributed technical teams.
• Proven track record in leading the initialization and deployment of large clusters or supercomputing environments.
• Experience with fast-paced AI labs or frontier-model infrastructure deployments and customer-facing architecture or field engineering roles.
• Strong systems engineering, coding, and debugging capabilities, including expertise in C/C++, Linux kernel, and drivers.
• Practical experience with NVIDIA GPU systems and SDKs such as CUDA, NVIDIA networking technologies, RoCE, InfiniBand, and/or ARM-based CPU solutions.
• Familiarity with virtualization and cloud-native networking principles.
• Willingness to travel up to 20% for on-site customer engagements and industry events.
• Equity
• Benefits
• Remote work locations
• Occasional travel opportunities for on-site customer visits and industry events
3Core Systems, Inc
Fortress Information Security
BizFirst LLC
Momentus Technologies
Get handpicked remote jobs straight to your inbox weekly.