
Manager, Solutions Architecture – AI Labs
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in California, +2 more states.
• Oversee the recruitment and management of a team consisting of solutions architects, systems/network engineers, and software engineers dedicated to large-scale GPU and AI networking implementations.
• Establish priorities, allocate resources effectively, mentor team members, and guarantee high-quality delivery for customers across multiple concurrent projects.
• Engage directly in technical reviews, design decisions, and essential debugging activities.
• Serve as the senior technical contact for a strategic AI lab, offering subject-matter expertise in advanced GPU and network systems.
• Direct compute and network configuration as well as performance troubleshooting to ensure the reliability of clusters.
• Facilitate discussions related to compute, network, and software architecture, while supporting server, network, and cluster deployment, which may include on-site data center work as necessary.
• Gather and synthesize specific customer requirements.
• Collaborate with GPU and Network Systems Engineering, Product Management, and Sales to shape roadmap priorities and develop package reference designs and solutions.
• Work alongside customer engineering and security teams to understand security needs and convert them into scalable AI infrastructure.
• Lead customer meetings, communicate project status and risks, and create design documents, debugging summaries, and presentations.
• Bachelor’s, Master’s, or PhD in Electrical/Computer Engineering, Computer Science, Physics, or a related Engineering field, or equivalent professional experience.
• At least 8 years of experience in Systems/Solutions/Field Engineering, Network or Data Center Engineering, or comparable roles.
• Minimum of 2 years in a leadership or mentoring capacity for engineers or architects.
• Direct experience in people management and recruitment for technical teams distributed across different geographical locations.
• In-depth knowledge of CPU/GPU server architecture, NICs, Linux, system software, and kernel drivers.
• Proficient in data center networking, including Ethernet and/or InfiniBand switches, NICs, fabrics, related tools, and cluster performance troubleshooting.
• Understanding of data center infrastructure, encompassing power, cooling, and deployment limitations.
• Capable of leading technical teams, setting priorities, and managing complex projects from design to production.
• Experience collaborating with Product Management, Sales, and Engineering teams.
• Strong time management skills, with the ability to balance strategic planning and hands-on support.
• Exceptional written and verbal communication skills, necessary for customer meetings, status updates, risk communication, design documentation, debugging summaries, and presentations.
• Proven history of leading the deployment and bring-up of large clusters or supercomputing environments.
• Background in AI labs or frontier-model infrastructure deployments in customer-facing roles.
• Proficient in systems engineering, coding, and debugging, including knowledge of C/C++, Linux kernel, and drivers.
• Practical experience with NVIDIA GPU systems and SDKs such as CUDA, NVIDIA networking technologies like NICs, RoCE, or InfiniBand, as well as ARM-based CPU solutions.
• Familiarity with virtualization and cloud-native networking principles.
• Willingness to travel occasionally, up to 20%, for on-site customer engagements and industry events.
• Equity
• Benefits
• Remote work options
• NVIDIA utilizes extensive conferencing tools
• Occasional travel up to 20% for on-site customer visits and industry events
3Core Systems, Inc
Fortress Information Security
BizFirst LLC
Momentus Technologies
Get handpicked remote jobs straight to your inbox weekly.