Principal Research Scientist, Synthetic Data Generation

Posted 14 hours ago

This is a fully remote position, open to applicants in California, +3 more states.

📋 Description

• Establish the technical vision for synthetic data generation in NVIDIA's frontier model initiatives

• Create and develop open-source libraries within the NVIDIA NeMo framework

• Produce synthetic datasets encompassing text, code, structured, and multimodal data for both pre- and post-training of LLMs

• Construct and enhance LLM-based data generation pipelines with automated quality assessment

• Innovate synthetic trajectories, multi-turn interactions, function calling, executable environments, reward modeling, and verifiable-reward data for training in agentic and tool-use contexts

• Propel advancements in multimodal synthetic data generation for images, documents, videos, and audio

• Promote safe and privacy-preserving synthesis through methods like differential privacy, anonymization, and de-identification

• Create and sustain open-source libraries and SDKs with well-defined APIs and comprehensive documentation

• Champion software excellence by utilizing modern tools, configurable architecture, and professional Git/CI-CD practices

• Contribute original research to leading conferences in machine learning and AI

• Collaborate with research, engineering, product, model teams, and external laboratories

• Guide and mentor scientists and engineers within the team


⛳️ Requirements

• PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience

• Over 15 years of engineering and research experience in synthetic data generation, generative modeling, multimodal machine learning, or related disciplines

• Profound technical knowledge of LLMs, pre-training, post-training, reinforcement learning stages, and inference frameworks such as vLLM or TGI

• Demonstrated success in developing or maintaining software libraries utilized by a wide developer community

• Experience in building and optimizing scalable data pipelines for large-scale model training, with a focus on throughput, distributed inference, and cost-effectiveness at cluster scale

• Strong publication history in top-tier venues such as NeurIPS, ICML, ICLR, ACL, or similar

• Notable open-source contributions in ML or data tooling, with significant community adoption

• Familiarity with multimodal generation or understanding, including vision-language, document AI, video, or audio

• Experience in generating data for agentic, tool-use, or reinforcement-learning post-training, including the design of RL environments

• Background in differential privacy, de-identification, or synthetic data applicable to regulated sectors like healthcare, finance, or government

• Experience in influencing model training decisions at the frontier scale, or direct collaboration with pre-training and post-training teams


🏝️ Benefits

• Equity

• Benefits

People also viewed

Center for Applied Linguistics15 hours ago

Research Assistant, Test Assembly, Desktop Publishing

US flagDistrict of Columbia, +1 more stateFull-timeResearch Scientist$61.3k/year
ApplyView job
CARE16 hours ago

Research Assistant

US flagUnited States OnlyFull-timeResearch Scientist$65k – $75k/year
ApplyView job
Sistema Fibra21 hours ago

Master's Research Fellow – Agent Experience Research

BR flagBrazil OnlyFull-timeResearch ScientistR$9,000/month
ApplyView job
Sistema Fibra21 hours ago

Master’s Research Fellow – Agent Experience Research

BR flagBrazil OnlyFull-timeResearch ScientistR$9,000/month
ApplyView job
JHU EEHPC3 days ago

Research Assistant – AI Development

US flagMaryland OnlyFull-timeResearch Scientist$17 – $31/hour
ApplyView job
Bureau Veritas Group3 days ago

Planning & Zoning Research Assistant

US flagUnited States OnlyPart-timeResearch Scientist$25 – $28/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers