
Senior Data & Document Ingestion Engineer β OCR, RAG
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in Poland, +5 more countries.
β’ Create and develop scalable pipelines for the ingestion of documents including PDFs, scans, emails, and office files.
β’ Integrate and enhance OCR and document extraction technologies to achieve high-accuracy text and layout extraction.
β’ Construct workflows for text cleansing, normalization, semantic chunking, and metadata tagging.
β’ Handle unstructured formats such as PDF, Word, Excel, and PowerPoint.
β’ Develop connectors for enterprise sources like SharePoint and email systems.
β’ Design data schemas and retrieval methods for downstream AI and RAG applications.
β’ Establish validation and monitoring loops to identify low-confidence OCR or extraction results.
β’ Ensure that ingestion pipelines comply with enterprise security, reliability, and latency standards.
β’ Implement logging, testing, and operational monitoring within data-processing workflows.
β’ Utilize Git, CI/CD, and best practices in software engineering for pipeline development.
β’ Between 5 to 10 years of professional experience in data engineering or backend/data-platform roles.
β’ Strong practical experience with Python and SQL.
β’ Demonstrated experience in building data ingestion and document-processing pipelines.
β’ Hands-on experience with unstructured documents, including PDF, Word, Excel, PPT, scans, or emails.
β’ Familiarity with OCR/document extraction tools such as AWS Textract or similar.
β’ Experience in constructing data-processing pipelines on public cloud platforms.
β’ Knowledge of AWS services like S3, Step Functions, and CloudWatch, or equivalent cloud services.
β’ Strong software development practices including Git, CI/CD, and automated testing.
β’ Experience with Azure, AWS, or Databricks in enterprise data environments is preferred.
β’ Preferred familiarity with vector databases, embeddings, or RAG architectures.
β’ Preferred experience in designing connectors for SharePoint, email, or other enterprise content systems.
β’ A background in insurance, financial services, or regulated-data environments is preferred.
β’ Proficiency in English is required.
β’ Competitive salary and performance-based bonuses.
β’ Flexible working hours and remote work options.
β’ Opportunities for professional development and continuous learning.
β’ Comprehensive health and wellness benefits.
β’ Collaborative and inclusive company culture.
Anyone AI
DriveNets
ServiceTitan
Get handpicked remote jobs straight to your inbox weekly.