
Principal Speech Data Linguist
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
β’ Take ownership of the linguistic framework for Innodata's segmentation and transcription tasks across various languages, domains, and applications.
β’ Establish standards for transcription and segmentation, including style guides and annotation conventions.
β’ Implement and manage quality frameworks, such as rubrics, error taxonomies, adjudication processes, inter-annotator agreement, and large-scale human quality assurance.
β’ Oversee the complete quality lifecycle for transcription and segmentation outputs.
β’ Create workflows that involve human oversight for review, correction, and adjudication as the quality of Automatic Speech Recognition (ASR) progresses.
β’ Address challenging cases involving accented and dialectal speech, low-resource and multilingual audio, overlapping speech, domain-specific jargon, and challenging acoustic environments.
β’ Collaborate with the Speech & Audio Research Scientist to convert model objectives into actionable specifications and evaluate the impact of transcription methods on ASR, Text-to-Speech (TTS), and speech/content-understanding models.
β’ Train, calibrate, and mentor expert transcribers and reviewers.
β’ Represent Innodata's approach to transcription and segmentation to clients and leading research labs.
β’ Contribute to the development of methodologies and best-practice documentation.
β’ Significant industry experience (typically over 8 years) in transcription, segmentation, and speech-data quality assurance.
β’ A Bachelor's degree in linguistics, phonetics, computational linguistics, or a closely related discipline is mandatory.
β’ Strong knowledge base in phonetics, phonology, and sociolinguistics.
β’ Comprehensive understanding of how choices in transcription and segmentation influence speech and content-understanding models.
β’ Proficiency in phonetic transcription and the International Phonetic Alphabet (IPA).
β’ Practical experience with acoustic and phonetic analysis, including spectrograms, formants, pitch, prosody, and segment boundaries.
β’ Extensive experience in audio segmentation, utterance and turn boundaries, timestamping, speaker labeling, and diarization.
β’ Proficient with tools such as Whisper, commercial ASR engines like AssemblyAI, Deepgram, Rev, and Speechmatics, Montreal Forced Aligner, and ELAN.
β’ Competency in Python scripting for batch processing, quality assurance, and metrics such as inter-annotator agreement and Word Error Rate (WER).
β’ Familiarity with regular expressions and Praat scripting.
β’ Solid data-management skills encompassing preprocessing, quality checks, post-processing, validation, and report preparation.
β’ Multilingual proficiency and hands-on experience with accented, dialectal, and code-switched speech.
β’ Excellent written and verbal communication skills.
β’ Comfortable collaborating with research scientists and clients.
β’ Bonus: Awareness of responsible AI considerations in speech, including accent and dialect bias, privacy, and consent regarding voice data.
β’ Competitive salary and performance-based bonuses.
β’ Opportunities for professional growth and development.
β’ Comprehensive health and wellness benefits.
β’ Flexible work arrangements, including remote work options.
β’ Collaborative and innovative work environment.
Orbital Engineering, Inc.
TRIGA Consulting GmbH & Co. KG
Curana Health
SEGUROS INBURSA S.A., GRUPO FINANCIERO INBURSA
Get handpicked remote jobs straight to your inbox weekly.