Senior Applied Research Scientist, Data Curation

Posted Aug 28

This is a fully remote position, open to applicants in California, +4 more states.

📋 Description

• Create efficient and high-performing models and data pipelines that extract and curate multimodal data—documents, images, audio, and videos—for foundation-model training.

• Construct petabyte-scale extraction and content-deduplication pipelines, encompassing document and HTML parsing, fuzzy and near-duplicate deduplication, semantic deduplication, and substring deduplication.

• Enhance and optimize curation techniques for petabyte-scale multimodal data across hundred-node GPU clusters.

• Investigate and design datasets, metrics, experiments, and validation scripts to establish standard research methodologies.

• Assist ML Engineers in scaling pipelines to production using NVIDIA Inference Microservices (NIMs) and blueprints.

• Author papers, blog posts, documentation, and training resources for customers.

• Stay informed on data-curation advancements in both academia and industry.

• Collaborate with Applied Research Scientists, Machine Learning Engineers, and MLOps Engineers.

• Guide junior engineers and interns.


⛳️ Requirements

• Master's, Ph.D., or equivalent experience in data curation, document AI, information retrieval, or multimodal research.

• Proven record of publications in prominent conferences such as CVPR, ICCV, ECCV, KDD, etc.

• Practical experience in developing computer vision and document-extraction models and pipelines, including layout analysis, OCR, and extraction of tables, figures, or formulas.

• Kaggle Grandmaster status or a strong track record of high-level results in machine learning competitions is highly desirable.

• In-depth understanding of the latest advancements in data curation research, particularly in multimodal content extraction and deduplication.

• Over 10 years of experience in developing multimodal systems across various models and platforms.

• Experience in information retrieval is a significant advantage.

• Proficiency in managing distributed data frameworks such as Ray, Spark, or Dask.

• Proven history of deploying large-scale, multi-node machine learning or data processing tasks in production environments.

• Familiarity with batching, streaming, and scaling ingestion pipelines.

• Excellent programming skills in Python.

• Strong practical experience with PyTorch or other modern deep learning frameworks.

• Ability to convey ideas through blog posts, papers, kernels, GitHub, etc.

• Outstanding communication and interpersonal skills.

• Capability to thrive in a dynamic, user-oriented, distributed team environment.


🏝️ Benefits

• Highly competitive salary.

• Equity.

• Comprehensive benefits package.

People also viewed

Mercor1 day ago

LLM Research Scientist – Pre-training, Computer Vision, Adversarial Robustness

US flagUnited States OnlyFreelanceResearch Scientist$100 – $120/hour
ApplyView job
Mercor1 day ago

LLM Research Scientist – Pre-training, Computer Vision, Adversarial Robustness

US flagUnited States OnlyFreelanceResearch Scientist$100 – $120/hour
ApplyView job
WestEd3 days ago

Research Assistant – Center for Mobility & Measurement

US flagCalifornia OnlyFull-timeResearch Scientist$71.2k – $111.3k/year
ApplyView job
Sistema Fibra3 days ago

Doctoral Research Fellow – Computer Science, Physics, Cybersecurity, Software Engineering

BR flagBrazil OnlyFull-timeResearch ScientistR$11k/month
ApplyView job
Sistema Fibra3 days ago

PhD Research Fellow – Computer Science, Physics, Cybersecurity, Software Engineering

BR flagBrazil OnlyFull-timeResearch ScientistR$11k/month
ApplyView job
Boston Medical Center (BMC)4 days ago

Research Assistant, Pediatrics

US flagMassachusetts OnlyFull-timeResearch Scientist$16 – $22/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers