
Senior Solutions Architect β Multimodal AI
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in Poland, +4 more countries.
β’ Establish technical connections with clients to develop multimodal AI systems aimed at document intelligence, personalization, and image/video analysis, encompassing everything from architectural design to production rollout.
β’ Advise clients on model training approaches across different modalities, such as layout-aware document encoders, user-item interaction models, and spatiotemporal video representations.
β’ Tackle visual content challenges, including issues related to image resolutions, optimizing vision encoders to meet production latency requirements, efficient video frame sampling, and temporal reasoning.
β’ Advocate for customer requirements to NVIDIA product teams, converting field experiences into strategic decisions for the roadmap concerning NeMo, TensorRT-LLM, Dynamo, and RAPIDS.
β’ Involve the developer community by organizing hackathons, delivering technical presentations, conducting demonstrations, and creating reference blueprints.
β’ Convert cutting-edge AI technologies into precise, production-ready solutions that provide measurable business advantages.
β’ MS or PhD in Computer Science, Engineering, or a related field will be considered.
β’ Over 7 years of experience in applied AI/ML.
β’ Practical experience in document understanding and visual content analysis.
β’ A proven history of constructing or enhancing VLMs and Omni models.
β’ Familiarity with NVIDIA's ecosystem, including TensorRT-LLM, NeMo, RAPIDS, or similar training and inference frameworks.
β’ Exceptional communication abilities; comfortable collaborating with research scientists, ML engineers, and business partners.
β’ Experience with multimodal AI systems that manage various modalities, including audio and video, and complex structures such as layouts, tables, and multi-page reasoning.
β’ Knowledge of retrieval and search systems, including dense retrieval, ANN indexing, and re-ranking pipelines.
β’ Experience in optimizing vision encoders for production through quantization, pruning, or architectural modifications.
β’ Published works or contributions to open-source projects in the field of multimodal learning.
β’ Highly competitive salaries
β’ Comprehensive benefits package
β’ Equal opportunity employment
U.S. Bank
First Quality
Arista Networks
Miovision
Get handpicked remote jobs straight to your inbox weekly.