
Senior Python Data Engineer – OCR & Document Processing
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Portugal, +1 more country.
• Design, construct, and enhance scalable data ingestion and document processing solutions for substantial volumes of unstructured insurance data.
• Develop and implement scalable pipelines for handling PDFs, scans, emails, and Office documents.
• Integrate, configure, and optimize OCR and document extraction technologies.
• Create automated workflows for document parsing, text cleaning, normalization, semantic chunking, and metadata enrichment.
• Build connectors and integrations for SharePoint, email systems, and enterprise repositories.
• Design and maintain vector database schemas and retrieval mechanisms for RAG solutions and AI applications.
• Ensure that pipelines comply with enterprise security, performance, compliance, and availability standards.
• Implement monitoring, validation, and quality-control mechanisms for low-confidence OCR and extraction results.
• Optimize workflows for scalability, reliability, and low-latency operations.
• Collaborate with AI Engineers, Backend Engineers, and Platform teams on comprehensive AI-powered document processing solutions.
• Develop and maintain cloud-native data ingestion solutions on public cloud platforms.
• 5-10 years of experience in Data Engineering, Data Processing, Document Intelligence, or related areas.
• Proven track record in building scalable data ingestion and processing pipelines.
• Experience managing large volumes of unstructured and semi-structured data.
• Proficiency in designing cloud-based data solutions.
• Strong programming abilities in Python.
• Solid knowledge of SQL.
• Practical experience with AWS services, including S3, Step Functions, and CloudWatch.
• Experience in processing unstructured documents such as PDF, Word, Excel, PowerPoint, and email content.
• Familiarity with building connectors and integrations with enterprise content repositories, such as SharePoint.
• Experience with OCR and document extraction tools, such as AWS Textract or comparable technologies.
• Experience in designing and implementing data ingestion and transformation pipelines.
• Understanding of vector databases and Retrieval-Augmented Generation (RAG) concepts.
• Familiarity with Git, CI/CD, and automated testing practices.
• Experience with Vector Databases.
• Exposure to RAG architectures and AI/LLM-based applications.
• Experience with Azure cloud services.
• Familiarity with Databricks.
• Experience in regulated industries such as Insurance or Banking.
• Full access to a foreign language learning platform.
• Personalized access to technology learning platforms.
• Customized workshops and training sessions to support your growth.
• Medical insurance coverage.
• Meal vouchers.
• Monthly budget for allocation on a flexible benefits platform.
• Access to 7 Card services.
• Wellbeing activities and social gatherings.
Mirantis
Solvd, Inc.
Loopio
Get handpicked remote jobs straight to your inbox weekly.