
Research Scientist, Video & Multimodal
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Collaborate directly with customers and frontier laboratories to develop models focused on video understanding, video-language, and video generation.
• Establish the design, structure, and evaluation methods for video data applicable to video and multimodal models.
• Convert model requirements into detailed data specifications that encompass modalities, annotation schemas, sampling techniques, and evaluation criteria.
• Create evaluation frameworks for video understanding that include temporal grounding, long-context reasoning, multi-turn interactions, cross-modal assessments, and retrieval-and-grounding evaluations.
• Develop evaluation methodologies for video generation that assess fidelity, temporal coherence, and physical plausibility.
• Determine the organization, enrichment, and sampling of video content, including complex, domain-specific footage.
• Conduct experiments and ablation studies linking data selection to quantifiable improvements in models.
• Design adversarial and challenging evaluations to detect system shortcomings and enhance data quality.
• Publish benchmarks, methodologies, and research papers.
• Collaborate with annotation teams, subject-matter experts, and synthetic data pipelines on data collection and labeling strategies.
• Approximately 5+ years of practical experience in the fields of video understanding or multimodal machine learning.
• A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative discipline is essential.
• An advanced degree (MS or PhD) in a relevant area is preferred.
• Experience in training and evaluating video or multimodal models, with a solid foundation in PyTorch.
• Proficiency in ffmpeg and decord pipelines, temporal and COCO-style annotations, WebDataset, Parquet and Arrow, as well as HuggingFace datasets.
• Experience in fine-tuning large video or vision-language models utilizing HuggingFace transformers, PEFT, and efficient inference methods.
• Familiarity with long-form video, streaming, temporal segmentation, or synthetic video generation.
• Experience in creating evaluation sets, calibrating difficulty levels, and defining quality video data for specified objectives.
• First-author publications or significant open-source contributions at conferences such as CVPR, ICCV, ECCV, NeurIPS, or ICLR.
• Ability to collaborate with research scientists and communicate data and modeling decisions effectively to both expert and non-expert audiences.
• A meticulous and reproducible approach to experimentation and documentation.
• Interest or practical experience in responsible AI evaluation and red-teaming is an added advantage.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and remote work options.
• Opportunities for professional development and continued education.
• A collaborative and innovative work environment.
Surgo Health
National University
University of Pennsylvania Perelman School of Medicine
Snorkel AI
Get handpicked remote jobs straight to your inbox weekly.