
Senior Data Engineer
Posted Jun 11

Posted Jun 11
This is a fully remote position, open to applicants in Brazil.
β’ You will design and enhance the datalake, which serves as the company's data backbone β the core system that supports, in real time, the dynamic pricing engine, machine learning models, and the group's business intelligence.
β’ This position entails ownership: you will establish the multi-tenant Lakehouse architecture, covering aspects from streaming to the semantic layer, while ensuring its reliability, governance, and cost-effectiveness.
β’ Develop and improve the data lake utilizing Apache Iceberg over S3 β implementing well-defined layers, partitioning and compaction, time-travel capabilities, and support for DELETE/UPDATE in accordance with LGPD (Brazilian data protection law).
β’ Create real-time ingestion processes (Kafka, Flink, CDC with Debezium) with managed schema evolution (Schema Registry) and delivery assurances.
β’ Design the transformation layer in dbt and coordinate batch and quality workflows in Airflow, spanning from crawler to backfill.
β’ Uphold metric definitions in Cube.js β the unified source that powers BI and AI agents, ensuring consistency throughout the organization.
β’ Execute federated and low-latency OLAP queries over the lake, maintaining cost and access isolation by tenant while ensuring high-performance queries.
β’ Guarantee data testing, lineage tracking, and cost efficiency, ensuring the platform remains reliable as it scales.
β’ Proficient in SQL with expertise in query optimization within distributed environments (Minimum 5 years).
β’ Experience in Python, particularly with PySpark or distributed processing.
β’ Knowledge of orchestration (Airflow), ELT processes, and dbt implemented at scale (Minimum 4 years).
β’ Familiarity with streaming technologies (Kafka, Flink) and Lakehouse architectures utilizing Apache Iceberg (Minimum 3 years).
β’ Strong grasp of data governance, quality assurance, and data modeling practices.
β’ Comfortable engaging with AI-assisted development tools (e.g., Claude Code).
β’ Experience with CDC (Debezium) and low-latency OLAP systems (ClickHouse, Pinot, Trino/Athena).
β’ Knowledge of semantic layers (Cube.js, dbt) and Data Mesh architectures.
β’ Familiarity with governance and cataloging tools (OpenMetadata, Lake Formation).
β’ Experience with vector databases (Qdrant) and data pipelines for machine learning.
β’ Remote work
β’ Project duration: 6 months, with the potential for extension or conversion to permanent employment.
Omada Health
BPO Global Services S.A.S
Get handpicked remote jobs straight to your inbox weekly.