
AI Research Engineer – Multi-Modal, Vision
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Switzerland.
• Perform comprehensive research and engineering on vision-language models, encompassing training, assessment, and optimization throughout the entire model development lifecycle.
• Create and execute post-training pipelines that include supervised fine-tuning, knowledge distillation, and reinforcement learning informed by human feedback.
• Develop and manage high-quality multimodal datasets, involving data curation, filtering, and balancing tailored for domain-specific tasks.
• Enhance model efficiency and deployability, modifying models for environments with limited resources using compression and optimization strategies.
• Design and establish evaluation frameworks and benchmarks to assess model performance, resilience, and success in real-world tasks.
• Construct and expand training workflows across a distributed GPU infrastructure.
• Identify and address bottlenecks in training pipelines to attain state-of-the-art model quality on specified benchmarks.
• Contribute to and utilize open-source ecosystems, including models, datasets, and tools, to expedite development.
• Keep abreast of the latest advancements in multimodal learning and vision-language systems, applying pertinent findings to drive practical enhancements.
• Publish research results in prestigious AI conferences and journals when applicable.
• Bachelor’s degree in Computer Science, Machine Learning, or a related discipline; MS/PhD is preferred.
• Extensive experience with multimodal post-training workflows, including supervised fine-tuning, knowledge distillation, and reinforcement learning based on feedback.
• Practical experience with parameter-efficient fine-tuning and distributed training frameworks.
• Proven capability to develop and enhance vision-language models with quantifiable results on standard benchmarks or real-world applications.
• Experience in adapting models for resource-limited environments.
• Established contributions to open-source projects in multimodal AI on platforms like GitHub or HuggingFace.
• Publications in leading AI conferences (NeurIPS, ICML, ICLR, CVPR, ECCV, etc.).
• Flexible work arrangements
• Professional development opportunities
Cerence Inc.
Tether.to
Cotiviti
Next Step Systems - National IT Recruiting Firm
Get handpicked remote jobs straight to your inbox weekly.