
Principal Research Scientist, Synthetic Data Generation
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in California, +3 more states.
• Establish the technical vision for synthetic data generation in NVIDIA's frontier model initiatives
• Create and develop open-source libraries within the NVIDIA NeMo framework
• Produce synthetic datasets encompassing text, code, structured, and multimodal data for both pre- and post-training of LLMs
• Construct and enhance LLM-based data generation pipelines with automated quality assessment
• Innovate synthetic trajectories, multi-turn interactions, function calling, executable environments, reward modeling, and verifiable-reward data for training in agentic and tool-use contexts
• Propel advancements in multimodal synthetic data generation for images, documents, videos, and audio
• Promote safe and privacy-preserving synthesis through methods like differential privacy, anonymization, and de-identification
• Create and sustain open-source libraries and SDKs with well-defined APIs and comprehensive documentation
• Champion software excellence by utilizing modern tools, configurable architecture, and professional Git/CI-CD practices
• Contribute original research to leading conferences in machine learning and AI
• Collaborate with research, engineering, product, model teams, and external laboratories
• Guide and mentor scientists and engineers within the team
• PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience
• Over 15 years of engineering and research experience in synthetic data generation, generative modeling, multimodal machine learning, or related disciplines
• Profound technical knowledge of LLMs, pre-training, post-training, reinforcement learning stages, and inference frameworks such as vLLM or TGI
• Demonstrated success in developing or maintaining software libraries utilized by a wide developer community
• Experience in building and optimizing scalable data pipelines for large-scale model training, with a focus on throughput, distributed inference, and cost-effectiveness at cluster scale
• Strong publication history in top-tier venues such as NeurIPS, ICML, ICLR, ACL, or similar
• Notable open-source contributions in ML or data tooling, with significant community adoption
• Familiarity with multimodal generation or understanding, including vision-language, document AI, video, or audio
• Experience in generating data for agentic, tool-use, or reinforcement-learning post-training, including the design of RL environments
• Background in differential privacy, de-identification, or synthetic data applicable to regulated sectors like healthcare, finance, or government
• Experience in influencing model training decisions at the frontier scale, or direct collaboration with pre-training and post-training teams
• Equity
• Benefits
Center for Applied Linguistics
CARE
Sistema Fibra
Sistema Fibra
Get handpicked remote jobs straight to your inbox weekly.