Data Scientist – AI Document Understanding, Co-op

Posted 9 hours ago

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Implement and explore innovative transformer and generative AI solutions for Document Understanding tasks.

β€’ Work on Optical Character Recognition (OCR), handwriting recognition, transcription, Named Entity Recognition, Relation Extraction, Coreference Resolution, Summarization, and Knowledge Graphs.

β€’ Process a variety of genealogical and historical collections, including newspapers, city directories, family history books, and vital records.

β€’ Assess the performance of multi-modal models in zero-shot and few-shot learning contexts.

β€’ Collaborate with ML Ops and Data Science Engineers to deploy datasets, truth sets, models, and pipelines for cloud-based training and inference.

β€’ Present insights, deliverables, and proposed solutions to both technical and non-technical audiences, including teams, stakeholders, and executives.

β€’ Construct, train, and fine-tune AI models that extract and organize text and image data from historical and genealogical records.

β€’ Train, optimize, and deploy models that support product development, customer success, and content creation across the Family History domain.


⛳️ Requirements

β€’ Currently pursuing an advanced degree (Master's or PhD preferred) in Computer Science, Data Science, Statistics, Mathematics, Linguistics, Engineering, or a related quantitative field with a strong focus on data.

β€’ Specialization in generative AI & LLMs, embeddings, LoRA, QLoRA, vector databases, transformer models, and Natural Language Processing (NLP).

β€’ Proficiency in software development, including data structures, distributed model training, and inference optimizations.

β€’ Strong command of Python and associated tools and libraries, including transformer models, multi-modal models, and general NLP.

β€’ Familiarity with Hugging Face Transformers, agentic frameworks and workflows, LangChain, LangGraph, and NLTK.

β€’ Knowledge of cloud platforms and AI/ML services such as Google Gemini API, Vertex AI, AWS EC2, S3, SageMaker, Model Registry, and Bedrock is advantageous.

β€’ All job offers are contingent upon a background check that complies with applicable laws.


🏝️ Benefits

β€’ Location-flexible work approach: nearest office, home, or a hybrid of both, subject to location restrictions and role requirements.

β€’ Reasonable accommodations for qualified individuals with disabilities.

People also viewed

Syndio7 hours ago

Staff Data Scientist, Data Products

US flagUnited States OnlyFull-timeData Scientist$180k – $205k/year
ApplyView job
Coinbase7 hours ago

Senior Data Scientist, Product

GB flagUnited Kingdom OnlyFull-timeData ScientistΒ£122.4k – Β£136k/year
ApplyView job
Freeport-McMoRan9 hours ago

Summer Internship – MIS Data Science

US flagArizona, +1 more stateInternshipData Scientist$29 – $36/hour
ApplyView job
CVS Health10 hours ago

Principal Data Scientist

US flagIllinois, +1 more stateFull-timeData Scientist$144.2k – $288.4k/year
ApplyView job
BeOne Medicines10 hours ago

Project Data Manager

US flagUnited States OnlyFull-timeData Scientist$123.2k – $163.2k/year
ApplyView job
Barr Engineering Co.10 hours ago

Data Science Intern

US flagMinnesota OnlyFull-timeData Scientist$30 – $38/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers