
Solutions Architect
Posted Jul 29

Posted Jul 29
This is a fully remote position, open to applicants in California.
• Successfully navigate the technical evaluation process: Collaborate with Sales to assess opportunities, drive technical discovery, and define evaluations based on the customer's actual success criteria rather than a generic checklist.
• Develop cluster and workload architectures that align with business outcomes: including time-to-first-token, tokens/sec, MFU, cost per training run, and reliability targets.
• Manage POCs from start to finish — defining scope, benchmarks, success metrics, timelines, and ensuring stakeholder alignment.
• Transition customers from contract signing to their first successful production training run: including provisioning, environment setup, validation benchmarks, and the essential debugging that occurs in between.
• Act as the designated technical owner for your accounts post-launch. You will be responsible for conducting architecture reviews, capacity planning, performance and cost optimization, and technical workload assessments.
• Identify opportunities for expansion proactively: understand where customers are capacity-constrained, what is on their roadmap, and what is required to address it.
• Build the function: Develop the resources that the Solutions Architecture team utilizes, such as demo environments, benchmarking tools, reference architectures, onboarding documentation, evaluation playbooks, and competitive materials.
• Provide high-quality feedback to Product, Engineering, and Research teams — highlighting recurring gaps, competitive losses, and the actual requests from customers once they are in production.
• Over 5 years of experience in customer-facing technical roles, including more than 2 years in pre-sales (Solutions Engineer, Sales Engineer, Solutions Architect, or specialist SA).
• Direct experience in selling or supporting GPU compute at a neocloud or GPU provider, or as an AI/HPC specialist at a hyperscaler or NVIDIA.
• Proven ability to guide a customer from evaluation to production while remaining accountable for the results.
• In-depth knowledge of large-scale training and inference: distributed training frameworks, multi-node topologies, InfiniBand/RoCE, storage and checkpointing, and understanding the limitations at scale.
• Familiarity with Kubernetes and SLURM as the scheduling environments that customers typically utilize.
• Proficiency in Python sufficient to create benchmarks, prototypes, or API integrations independently instead of relying on engineering support.
• A history of managing technical evaluations in complex, multi-stakeholder situations and influencing the outcomes positively.
• Excellent communication skills, capable of engaging in detailed discussions with both distributed systems engineers and CFOs.
• Comprehensive benefits package for you and your dependents, which includes healthcare, dental, and vision coverage.
• 401(k) plan.
• Unlimited Paid Time Off (PTO).
Snowflake
Cisco
BCD Travel
Get handpicked remote jobs straight to your inbox weekly.