
Research Scientist, Speech & Audio
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Develop the design, structure, and assessment frameworks for audio data utilized in speech and audio models.
• Convert requirements for speech and audio models into detailed data specifications, including modalities, transcription and annotation schemas, sampling methods, and evaluation standards.
• Create evaluation methodologies that address semantic accuracy, robustness to noise and accents, code-switching, diarization error rates, speech naturalness and intelligibility, as well as streaming latency.
• Organize, enhance, and sample audio from diverse languages, accents, acoustic environments, demographic groups, emotional tones, paralinguistic ranges, scripted versus spontaneous speech, and various speaker configurations.
• Collaborate with transcription and linguistics leaders to formulate transcription specifications and assess their influence on model performance.
• Work alongside audio solutions and engineering teams to guarantee that collected audio aligns with model objectives.
• Conduct fine-tuning and evaluation experiments with ablation studies that correlate data selection with quantifiable improvements.
• Design adversarial tests and challenging evaluations to discover failures in speech systems.
• Publish benchmarks, methodologies, and research findings.
• Collaborate with annotation teams, subject-matter experts, and synthetic- and augmented-audio pipelines to implement specifications effectively.
• Engage directly with customers and leading labs that are developing ASR, text-to-speech, speech-to-speech, conversational voice, diarization, and audio-language models.
• Approximately 5+ years of practical experience in the speech or audio machine learning industry.
• A Bachelor's degree in computer science, electrical engineering, or a related technical or quantitative discipline is mandatory.
• Experience in training and evaluating speech or audio models — including ASR, TTS, speech-to-speech, speaker, or audio-language models — with a solid foundation in PyTorch.
• Proficiency in ESPnet, NeMo, SpeechBrain, Kaldi, HuggingFace, forced alignment, WER/CER, and other relevant speech metrics.
• Practical experience with multilingual, accented, dialectal, low-resource, or code-switched speech.
• Familiarity with synthetic or augmented audio, encompassing TTS pipelines, noise simulation, and room-response modeling.
• Experience in creating evaluation sets and analyzing dataset coverage across various conditions.
• First-author publications or significant contributions to open-source projects in venues such as Interspeech, ICASSP, ASRU, SLT, or NeurIPS.
• Capability to collaborate directly with research scientists from customer and frontier labs.
• Ability to clearly communicate data and modeling choices to both expert and lay audiences.
• A methodical and reproducible approach to experimentation and documentation.
• Interest or practical experience in responsible AI evaluation and red-teaming is an advantage.
• Competitive salary and performance-based incentives.
• Comprehensive health, dental, and vision insurance.
• Opportunities for professional development and career advancement.
• Flexible work hours and remote work options.
• Collaborative and innovative work environment.
Surgo Health
National University
University of Pennsylvania Perelman School of Medicine
Snorkel AI
Get handpicked remote jobs straight to your inbox weekly.