
Principal Solutions Architect
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in Maryland.
β’ Collaborate with NVIDIA Cloud Partners to design, implement, and deliver NVIDIA's cutting-edge hardware and software solutions.
β’ Work alongside SAs, Account Managers, Engineering, Product, and business leaders to synchronize strategies, evaluate technical requirements, and secure business opportunities for NVIDIA.
β’ Act as the primary technical resource for customers throughout the design, development, construction, integration, and production phases of GPU Cloud infrastructure and applications during the entire customer lifecycle.
β’ Conduct regular technical meetings with customers to discuss project/product specifics, feature enhancements, introduce new technologies, and facilitate debugging sessions.
β’ Collaborate closely with customers to develop and implement NVIDIA solutions, including Proofs of Concept (PoCs), to meet essential business requirements across infrastructure, libraries, and applications.
β’ Prepare and present technical content to customers, which includes presentations, workshops, reference architectures, tutorials, and publications.
β’ Promote usage and integration by incorporating libraries, frameworks, models, and software applications.
β’ Deliver GenAI, AI, and ML hardware/software to production alongside key customers and partners.
β’ Manage end-to-end technology solution integration with strategic customers and provide product strategy recommendations based on insights gathered.
β’ BS/MS/PhD in Electrical/Computer Engineering, Computer Science, Physics, Mathematics, or other Engineering disciplines, or equivalent experience.
β’ Over 15 years of experience in Solution Engineering (or similar roles such as Sales Engineering, Cloud Engineering, or Solution Architecture), including direct engagement with partners and customers.
β’ Proven experience in designing and deploying large-scale cluster environments.
β’ Hands-on experience in designing, developing, and delivering distributed Cloud architectures.
β’ Strong foundational knowledge in programming, optimizations, and software design, particularly in Python and Deep Learning frameworks like PyTorch and TensorFlow.
β’ Practical expertise in fine-tuning and deploying models, as well as integrating software application stacks, libraries, and frameworks to enhance usage of GPU platforms.
β’ A proactive approach and skills to oversee and drive complex multi-disciplinary technical engagements with customers throughout the complete customer lifecycle and across functional teams.
β’ Effective time management skills and the ability to juggle multiple tasks efficiently.
β’ Outstanding presentation, communication, and collaboration abilities.
β’ Self-motivated individual with a passion for growth, continuous learning, and sharing insights.
β’ Hands-on experience with NVIDIA GPUs, software libraries, frameworks, and foundational models such as NVIDIA Nemotron, NVIDIA NeMo Framework, NVIDIA Dynamo, NeMo Retriever, NVIDIA Triton Inference Server, TensorRT, TensorRT-LLM, and NVIDIA CUDA-X.
β’ Practical expertise with large-scale AI cloud environments (e.g., AWS, Azure, GCP) and on-premises/hybrid infrastructures, particularly for inference and training workloads.
β’ Familiarity with NVIDIA hardware (such as GPUs, networking, storage) and systems technologies including NCCL, DCGM, UFM, Mission Control, and Base Command Manager.
β’ Proficient in large-scale AI model training/deployment involving GPU systems, performance testing, AI benchmarking, fine-tuning, with a strong emphasis on MLOps and cluster orchestration (SLURM, K8s, orchestrator, load balancing, cloud architecture).
β’ Experience working with enterprise developers and strong customer-facing skills.
β’ Equity
β’ Benefits
CACI International Inc
QAD
jaydhub
Sitetracker
Get handpicked remote jobs straight to your inbox weekly.