
Data Engineer – Complex Data Pipelines
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in France.
• Create and develop data pipelines from the ground up, encompassing data ingestion, processing, transformation, storage, and consumption.
• Design data ingestion systems for distributed recording nodes that are frequently offline, incorporating features such as local buffering, resumable transfers, and the reconciliation of late-arriving or out-of-order data upon reconnection.
• Collaborate with the ML team to determine which data should be prioritized during short or bandwidth-constrained connection intervals.
• Create pipelines capable of processing structured/tabular data, images, video, and temporal/time-series data.
• Work with sequential data generated by sensors and manage clock drift across nodes to ensure reliable downstream timing.
• Build robust and scalable data processing solutions utilizing Python and SQL.
• Design data models and storage strategies, including capacity planning and retention for substantial volumes of image and video data on self-managed storage.
• Oversee workflow definitions in the orchestration layer, including aspects such as ordering, retry mechanisms, idempotency, and backfill behavior.
• Establish processes for data ingestion, transformation, validation, quality control, and traceability.
• Develop tools that facilitate data preparation and availability for machine learning and AI applications.
• Collaborate with Machine Learning Engineers, DevOps, and Software Engineers to comprehend data needs and deliver solutions.
• Ensure that pipelines remain reliable, maintainable, and scalable as data volumes and use cases evolve.
• Identify and address data-quality issues, including gaps and duplicates resulting from node outages and retries, and establish monitoring for data quality and pipeline health.
• Define the overarching architecture and technical standards for the in-house data platform.
• Document pipeline architecture, data flows, and technical solutions.
• Extensive professional experience in a closely related data engineering position.
• Proven experience in designing and implementing data pipelines from scratch (not cloud-based), including making architectural and technical decisions.
• Proficient programming skills in Python and SQL.
• Professional experience handling various types of data.
• Experience constructing pipelines that involve at least some of the following: images, video, sensor data, time-series, or other sequential data.
• Familiarity with systems that must accommodate unreliable or absent network connectivity and recover smoothly, including offline-first, store-and-forward, edge collection, or similar architectures.
• Experience managing data infrastructure on bare metal or self-managed servers, rather than solely on managed cloud services.
• Strong understanding of data ingestion, transformation, storage, validation, and data-quality principles.
• Experience working with large or complex datasets.
• Solid Linux skills, including proficiency with filesystems, storage, services, and network troubleshooting.
• Good software engineering practices, including experience with Git, testing, code review, and documentation.
• Ability to independently investigate technical issues and propose suitable architecture and solutions.
• Fluent in English, both written and spoken.
• Full-time employment (CDI)
Imagemaker
Vidmob
Get handpicked remote jobs straight to your inbox weekly.