Remotery

Senior Data Engineer

Posted Jun 11

This is a fully remote position, open to applicants in Brazil.

πŸ“‹ Description

β€’ You will design and enhance the datalake, which serves as the company's data backbone β€” the core system that supports, in real time, the dynamic pricing engine, machine learning models, and the group's business intelligence.

β€’ This position entails ownership: you will establish the multi-tenant Lakehouse architecture, covering aspects from streaming to the semantic layer, while ensuring its reliability, governance, and cost-effectiveness.

β€’ Develop and improve the data lake utilizing Apache Iceberg over S3 β€” implementing well-defined layers, partitioning and compaction, time-travel capabilities, and support for DELETE/UPDATE in accordance with LGPD (Brazilian data protection law).

β€’ Create real-time ingestion processes (Kafka, Flink, CDC with Debezium) with managed schema evolution (Schema Registry) and delivery assurances.

β€’ Design the transformation layer in dbt and coordinate batch and quality workflows in Airflow, spanning from crawler to backfill.

β€’ Uphold metric definitions in Cube.js β€” the unified source that powers BI and AI agents, ensuring consistency throughout the organization.

β€’ Execute federated and low-latency OLAP queries over the lake, maintaining cost and access isolation by tenant while ensuring high-performance queries.

β€’ Guarantee data testing, lineage tracking, and cost efficiency, ensuring the platform remains reliable as it scales.


⛳️ Requirements

β€’ Proficient in SQL with expertise in query optimization within distributed environments (Minimum 5 years).

β€’ Experience in Python, particularly with PySpark or distributed processing.

β€’ Knowledge of orchestration (Airflow), ELT processes, and dbt implemented at scale (Minimum 4 years).

β€’ Familiarity with streaming technologies (Kafka, Flink) and Lakehouse architectures utilizing Apache Iceberg (Minimum 3 years).

β€’ Strong grasp of data governance, quality assurance, and data modeling practices.

β€’ Comfortable engaging with AI-assisted development tools (e.g., Claude Code).

β€’ Experience with CDC (Debezium) and low-latency OLAP systems (ClickHouse, Pinot, Trino/Athena).

β€’ Knowledge of semantic layers (Cube.js, dbt) and Data Mesh architectures.

β€’ Familiarity with governance and cataloging tools (OpenMetadata, Lake Formation).

β€’ Experience with vector databases (Qdrant) and data pipelines for machine learning.


🏝️ Benefits

β€’ Remote work

β€’ Project duration: 6 months, with the potential for extension or conversion to permanent employment.

People also viewed

RemofirstJul 26

Senior Data Engineer

EG flagEgypt OnlyFull-timeData Engineer
ApplyView job
Omada HealthJul 26

Staff Software Engineer, Data Products

US flagUnited States OnlyFull-timeData Engineer$202.4k – $253k/year
ApplyView job
MoovxJul 26

Senior Data Engineer

Latin AmericaFull-timeData Engineer
ApplyView job
BPO Global Services S.A.SJul 25

Data Engineer

CO flagColombia OnlyFull-timeData Engineer$10/hour
ApplyView job
GSB SolutionsJul 25

Technical Program Manager – Data & Power Platform

MX flagMexico OnlyFull-timeData Engineer$113k/year
ApplyView job
DOMVS iTJul 25

Data Engineer, Mid/Senior

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers