
Speech Research Intern β 2
Posted Jul 15

Posted Jul 15
This is a fully remote position, open to applicants in United States.
β’ Create and assess speech-first models, with an emphasis on Spoken Language Models (SLMs)
β’ Transition concepts from prototypes to functional demonstrations
β’ Collaborate with scientists and engineers to achieve measurable outcomes
β’ Build comprehensive speech dialogue systems and speech-aware large language models (LLMs)
β’ Maintain consistency between speech encoders and text backbones through lightweight adapters
β’ Perform reliable evaluations across tasks related to recognition, understanding, and generation
β’ Implement latency-aware inference to enhance real-time user interactions
β’ PhD candidate in Computer Science/Electrical Engineering (or related field) with research experience in speech, audio machine learning, or multimodal language models
β’ Proficiency in Python and PyTorch, along with practical experience in GPU training
β’ Familiarity with torchaudio or librosa
β’ Solid understanding of contemporary sequence models (Transformers or State Space Models) and best practices for training
β’ Expertise in at least one specific area: (a) discrete speech tokens or temporal compression, (b) modality alignment to LLMs using adapters, or (c) post-training or instruction tuning for speech-related tasks
β’ Strong experimental skills: clean coding practices, ablation studies, reproducibility, and effective reporting
β’ Competitive stipend
β’ Guidance from applied scientists and engineers
β’ Opportunities to publish and present research
β’ Access to state-of-the-art GPU infrastructure
β’ Supportive setting for rapid and responsible experimentation
β’ Flexible options for location and scheduling
Julesetmoi
National University
MeridianLink
Get handpicked remote jobs straight to your inbox weekly.