
AI Research Engineer – Multi-Modal, Vision
Posted Jun 17

Posted Jun 17
This is a fully remote position, open to applicants in United Arab Emirates (UAE).
• Conduct comprehensive research and engineering on vision-language models, encompassing training, evaluation, and optimization throughout the entire model development lifecycle.
• Design and execute post-training pipelines that include supervised fine-tuning, knowledge distillation, and reinforcement learning guided by human feedback.
• Develop and sustain high-quality multimodal datasets, involving data curation, filtering, and balancing tailored for specific domain tasks.
• Enhance model efficiency and deployability by adapting models for environments with limited resources through compression and optimization techniques.
• Create and implement evaluation frameworks and benchmarks aimed at assessing model performance, robustness, and success in real-world applications.
• Construct and scale training workflows utilizing distributed GPU infrastructure.
• Identify and address bottlenecks in training pipelines to attain state-of-the-art model quality on designated benchmarks.
• Contribute to and utilize open-source ecosystems, including models, datasets, and tools, to expedite development.
• Remain updated on the latest advancements in multimodal learning and vision-language systems, applying relevant insights to enhance practical applications.
• Publish research outcomes in leading AI conferences and journals whenever applicable.
• A degree in Computer Science, Machine Learning, or a related discipline; an MS/PhD is preferred.
• Extensive experience with multimodal post-training workflows, including supervised fine-tuning, knowledge distillation, and reinforcement learning based on feedback.
• Practical experience with parameter-efficient fine-tuning and distributed training frameworks.
• Proven capability to develop and refine vision-language models, demonstrating measurable success on standard benchmarks or real-world tasks.
• Experience in adapting models for environments with limited resources.
• Established open-source contributions in multimodal AI on platforms like GitHub or HuggingFace.
• Publications in top-tier AI conferences such as NeurIPS, ICML, ICLR, CVPR, ECCV, etc.
• Health insurance
• Flexible work arrangements
• Professional development opportunities
Cerence Inc.
Tether.to
Tether.to
Get handpicked remote jobs straight to your inbox weekly.