
Senior Data Engineer – Databricks, DBT
Posted Aug 27

Posted Aug 27
This is a fully remote position, open to applicants in Brazil.
• Develop and enhance the company's new data platform utilizing DBT on Databricks.
• Architect and implement the DBT platform structure on Databricks, incorporating Unity Catalog and medallion layers (bronze/silver/gold).
• Organize the DBT project according to modeling, naming, layering, macros, testing, and documentation standards.
• Utilize dimensional modeling techniques (Kimball, star schema) and/or Data Vault when required.
• Create CI/CD pipelines for DBT via GitHub Actions, encompassing validation, builds, testing, and deployment across development, staging, and production environments.
• Automate the provisioning and operations using Infrastructure as Code (IaC) with Terraform and Databricks Asset Bundles.
• Design and sustain Databricks Workflows for orchestrating data pipelines.
• Oversee reproducible environments with code versioning utilizing Git, GitFlow, or trunk-based development methodologies.
• Enhance the cost-efficiency and performance of clusters, SQL Warehouses, Photon, partitioning, and incremental modeling.
• Work with Delta Lake, Parquet, and various distributed data formats.
• Employ Apache Spark and PySpark for scenarios that extend beyond SQL when applicable.
• Establish data governance and security measures, including Unity Catalog, access control, and the masking of sensitive information.
• Guarantee adherence to Brazil’s LGPD and best practices in data privacy.
• Guide the team on best practices in analytics engineering.
• Evaluate code (PRs) to ensure the technical quality of outcomes.
• Collaborate in agile and cooperative environments.
• DBT Core: experience with incremental models, snapshots, macros (Jinja), tests, packages, exposures — at least 4 years.
• Databricks: familiarity with Unity Catalog, SQL Warehouses, clusters, Delta Lake, and Workflows — at least 4 years.
• Proficiency in Apache Spark and distributed data processing.
• Strong knowledge of Python and advanced SQL.
• Experience with GitHub Actions for data pipeline CI/CD.
• Knowledge of Databricks Asset Bundles for packaging and deployment.
• Competency in Git and collaborative workflows (trunk-based development or GitFlow, PRs, code review).
• Understanding of cloud computing, preferably GCP (Azure or AWS also considered).
• Experience with ETL/ELT and data integration tools.
• Commitment to a culture of automated testing, version control, and reproducible deployments.
• Experience with dbt Fusion / migration to dbt Core 2.0 — a plus.
• Familiarity with Terraform and IaC for managing environments and secrets — a plus.
• Experience with orchestration tools like Airflow or Dagster — preferred.
• Knowledge of Kubernetes and container technologies (Docker) — preferred.
• Experience in data observability tools: dbt docs, Elementary, Monte Carlo, OpenLineage — preferred.
• Background in sensitive data masking / LGPD and compliance — preferred.
• Expertise in Kimball dimensional modeling / star schema / Data Vault — preferred.
• Porto Seguro health insurance, with the option to add a spouse and children.
• Porto Seguro dental coverage for employees and their dependents.
• Profit-Sharing and Results Program (PLR).
• Childcare support.
• Alelo meal and food vouchers.
• Home office allowance.
• Collaborations with educational institutions, providing discounts and incentives for courses and degree programs.
• Certification incentives, including for cloud certifications (GCP, Azure, AWS, and others).
• Livelo points.
• TotalPass, offering discounted gym memberships for employees and their family members.
• Mindself, providing incentives for meditation and mindfulness practices.
• Fully remote work opportunities.
• Continuous professional development initiatives.
Mirantis
Solvd, Inc.
Loopio
Get handpicked remote jobs straight to your inbox weekly.