Senior Python Data Engineer – OCR & Document Processing

Posted 3 days ago

This is a fully remote position, open to applicants in Portugal, +1 more country.

📋 Description

• Design, construct, and enhance scalable data ingestion and document processing solutions for substantial volumes of unstructured insurance data.

• Develop and implement scalable pipelines for handling PDFs, scans, emails, and Office documents.

• Integrate, configure, and optimize OCR and document extraction technologies.

• Create automated workflows for document parsing, text cleaning, normalization, semantic chunking, and metadata enrichment.

• Build connectors and integrations for SharePoint, email systems, and enterprise repositories.

• Design and maintain vector database schemas and retrieval mechanisms for RAG solutions and AI applications.

• Ensure that pipelines comply with enterprise security, performance, compliance, and availability standards.

• Implement monitoring, validation, and quality-control mechanisms for low-confidence OCR and extraction results.

• Optimize workflows for scalability, reliability, and low-latency operations.

• Collaborate with AI Engineers, Backend Engineers, and Platform teams on comprehensive AI-powered document processing solutions.

• Develop and maintain cloud-native data ingestion solutions on public cloud platforms.


⛳️ Requirements

• 5-10 years of experience in Data Engineering, Data Processing, Document Intelligence, or related areas.

• Proven track record in building scalable data ingestion and processing pipelines.

• Experience managing large volumes of unstructured and semi-structured data.

• Proficiency in designing cloud-based data solutions.

• Strong programming abilities in Python.

• Solid knowledge of SQL.

• Practical experience with AWS services, including S3, Step Functions, and CloudWatch.

• Experience in processing unstructured documents such as PDF, Word, Excel, PowerPoint, and email content.

• Familiarity with building connectors and integrations with enterprise content repositories, such as SharePoint.

• Experience with OCR and document extraction tools, such as AWS Textract or comparable technologies.

• Experience in designing and implementing data ingestion and transformation pipelines.

• Understanding of vector databases and Retrieval-Augmented Generation (RAG) concepts.

• Familiarity with Git, CI/CD, and automated testing practices.

• Experience with Vector Databases.

• Exposure to RAG architectures and AI/LLM-based applications.

• Experience with Azure cloud services.

• Familiarity with Databricks.

• Experience in regulated industries such as Insurance or Banking.


🏝️ Benefits

• Full access to a foreign language learning platform.

• Personalized access to technology learning platforms.

• Customized workshops and training sessions to support your growth.

• Medical insurance coverage.

• Meal vouchers.

• Monthly budget for allocation on a flexible benefits platform.

• Access to 7 Card services.

• Wellbeing activities and social gatherings.

People also viewed

Fueled21 hours ago

Google Cloud Data Engineer

Latin AmericaFreelanceData Engineer
ApplyView job
Mirantis21 hours ago

Senior Data Platform Engineer – Kafka, PostgreSQL

KZ flagKazakhstan OnlyFull-timeData Engineer
ApplyView job
Solvd, Inc.22 hours ago

Snowflake Data Architect

IN flagIndia OnlyFull-timeData Engineer
ApplyView job
Loopio22 hours ago

Data Engineer

CA flagCanada OnlyFull-timeData EngineerC$103.5k – C$140k/year
ApplyView job
Blend36022 hours ago

Lead Data Engineer

AR flagArgentina OnlyFull-timeData Engineer
ApplyView job
SysMap Solutions23 hours ago

Mid-Level/Senior Data Engineer

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers