
Research Engineer – Data Infrastructure
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in United Kingdom.
• Develop extensive data pipelines for the collection, processing, filtering, and transformation of datasets utilized to train cutting-edge models.
• Train models that are integral to data processing pipelines, which include classifiers, quality filters, and labeling models.
• Design strategies for data curation such as deduplication, quality scoring, labeling, and augmentation to enhance model performance.
• Build tools and infrastructure that empower researchers to efficiently and reliably explore and train on vast datasets.
• Oversee the data infrastructure that drives ElevenLabs' leading AI models.
• No official certifications or degrees are necessary.
• Passionate engineers who can showcase their ability to tackle exceptionally challenging problems through previous projects, designs, or contributions on GitHub.
• Experience in developing data-intensive systems, preferably those that support machine learning training pipelines.
• Strong engineering capabilities in distributed data processing at scale, including technologies like Kubernetes or custom pipelines for large datasets.
• Ability to independently assess how data quality, composition, and curation influence model results, and to create tools for measurement.
• Experience in building or managing web crawlers is an advantage.
• Annual discretionary professional development stipend.
• Annual discretionary stipend for social travel to connect with colleagues.
• Annual company offsite.
• Monthly co-working stipend for employees located away from a main hub.
• Option to work from ElevenLabs offices in London, New York, San Francisco, and Warsaw.
Data Elephant
ICF
General Dynamics Information Technology
Logic20/20, Inc.
Get handpicked remote jobs straight to your inbox weekly.