
Senior Solutions Architect – Generative AI Deployment, AIOps
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in California, +4 more states.
• Collaborate with other solution architects, engineers, product teams, and business units to comprehend strategies and technical requirements, ultimately defining high-value solutions.
• Interact with developers, scientific researchers, and data scientists across various technical domains.
• Form strategic partnerships with key customers and industry-specific solution partners to leverage NVIDIA’s computing platform.
• Assist customers in adopting and developing innovative solutions utilizing NVIDIA technology and MLOps solutions.
• Evaluate the performance and power efficiency of AI inference workloads within Kubernetes.
• Act as a trusted technical consultant on projects and proof-of-concept initiatives focused on inference for Generative AI and Large Language Models.
• Work alongside internal teams to conduct performance analysis and modeling of inference software.
• Travel to conferences and customer sites as needed (20%).
• Bachelor’s, Master’s, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or other related fields (or equivalent experience).
• Over 8 years of hands-on experience with Deep Learning frameworks such as PyTorch and TensorFlow.
• Strong foundational skills in programming, optimization, and software design, particularly in Python.
• Proficient in problem-solving and debugging in GPU orchestration and Multi-Instance GPU (MIG) management within Kubernetes environments.
• Experience with containerization and orchestration technologies, as well as monitoring and observability solutions for AI deployments.
• Comprehensive understanding of the theory and practice surrounding LLM and DL inference.
• Exceptional presentation, communication, and collaboration abilities.
• Prior experience with large-scale DL training and deploying or optimizing DL inference in production environments.
• Familiarity with NVIDIA GPUs and software libraries such as NVIDIA NIM, Dynamo, TensorRT, and TensorRT-LLM.
• Strong C/C++ programming skills, including debugging, profiling, code optimization, performance analysis, and test design.
• Knowledge of parallel programming and distributed computing platforms.
• Some travel to conferences and customer locations may be necessary (20%).
• Equity
• Benefits
NVIDIA
Cisco
Gateway Ticketing Systems UK Ltd
MoneyGram
Get handpicked remote jobs straight to your inbox weekly.