
Senior Applied Research Scientist, Data Curation
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in California, +4 more states.
• Create efficient and high-performing models and data pipelines that extract and curate multimodal data—documents, images, audio, and videos—for foundation-model training.
• Construct petabyte-scale extraction and content-deduplication pipelines, encompassing document and HTML parsing, fuzzy and near-duplicate deduplication, semantic deduplication, and substring deduplication.
• Enhance and optimize curation techniques for petabyte-scale multimodal data across hundred-node GPU clusters.
• Investigate and design datasets, metrics, experiments, and validation scripts to establish standard research methodologies.
• Assist ML Engineers in scaling pipelines to production using NVIDIA Inference Microservices (NIMs) and blueprints.
• Author papers, blog posts, documentation, and training resources for customers.
• Stay informed on data-curation advancements in both academia and industry.
• Collaborate with Applied Research Scientists, Machine Learning Engineers, and MLOps Engineers.
• Guide junior engineers and interns.
• Master's, Ph.D., or equivalent experience in data curation, document AI, information retrieval, or multimodal research.
• Proven record of publications in prominent conferences such as CVPR, ICCV, ECCV, KDD, etc.
• Practical experience in developing computer vision and document-extraction models and pipelines, including layout analysis, OCR, and extraction of tables, figures, or formulas.
• Kaggle Grandmaster status or a strong track record of high-level results in machine learning competitions is highly desirable.
• In-depth understanding of the latest advancements in data curation research, particularly in multimodal content extraction and deduplication.
• Over 10 years of experience in developing multimodal systems across various models and platforms.
• Experience in information retrieval is a significant advantage.
• Proficiency in managing distributed data frameworks such as Ray, Spark, or Dask.
• Proven history of deploying large-scale, multi-node machine learning or data processing tasks in production environments.
• Familiarity with batching, streaming, and scaling ingestion pipelines.
• Excellent programming skills in Python.
• Strong practical experience with PyTorch or other modern deep learning frameworks.
• Ability to convey ideas through blog posts, papers, kernels, GitHub, etc.
• Outstanding communication and interpersonal skills.
• Capability to thrive in a dynamic, user-oriented, distributed team environment.
• Highly competitive salary.
• Equity.
• Comprehensive benefits package.
Mercor
Mercor
WestEd
Sistema Fibra
Get handpicked remote jobs straight to your inbox weekly.