
Senior Data Engineer – Data Architecture
Posted Sep 1

Posted Sep 1
This is a fully remote position, open to applicants in Brazil.
• Create, develop, and manage ingestion and transformation pipelines orchestrated through Apache Airflow, adhering to consistent standards for retries, idempotency, alerting, and SLAs.
• Build and enhance dbt transformation models within the data warehouse, structuring staging, intermediate, and marts layers with thorough tests and documentation.
• Facilitate ingestion from diverse sources, such as transactional databases, third-party APIs, regulatory files, and operational spreadsheets.
• Implement data observability practices, including monitoring freshness, volume, contract violations, and source-to-target reconciliation.
• Establish and promote the warehouse data model, covering keys, grain, historization (SCD), and the management of retroactive corrections.
• Set standards for distribution, sort-order, and partitioning within the MPP environment, ensuring compliance with these standards.
• Design data layers and contracts between teams, specifying the source of truth, write permissions, and accessibility for BI.
• Oversee performance and cost optimization, including plan analysis, queues/WLM, concurrency, maintenance, query rewrites, and materializations.
• Assess and propose technologies based on clear trade-offs, costs, and potential exit costs.
• Establish automated quality controls and reconciliations for customer reporting and regulatory compliance.
• Guarantee end-to-end lineage and traceability, from source and code version to the figures represented in dashboards.
• Implement access controls, sensitivity-based segregation, and adherence to LGPD, CVM, BACEN, BSM, and internal compliance policies.
• Maintain documentation of architectural decisions (ADRs) and keep the data dictionary current.
• Completed Bachelor’s degree.
• Proficient in advanced SQL: window functions, recursive CTEs, execution plan analysis, and diagnosing data skew and disk spills.
• Experience with Python for data engineering: modular, testable, version-controlled code—not limited to notebooks.
• Familiarity with pandas/Polars, typing, and automated testing practices.
• Hands-on experience with Apache Airflow in production: DAG authoring, sensors, backfills, dependency management, and failure handling.
• Experience with cloud-based MPP data warehouses—Amazon Redshift preferred; Snowflake, BigQuery, or Databricks accepted, with a readiness to deepen expertise in Redshift.
• Proficient in dbt or a similar version-controlled transformation tool with testing capabilities.
• Knowledge of dimensional modeling (Kimball) with practical expertise in fact vs. dimension, grain, Type 1/Type 2 SCDs, bridge tables, and snapshots.
• Familiar with Git and a culture of code review; CI/CD principles applied to data.
• Capable of performance modeling and diagnosis: able to articulate the reasons behind slow queries and demonstrate the improvements with before-and-after metrics.
• Experience in the financial services sector is an advantage.
• 20 days of planned leave.
• TotalPass.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.