
Senior Data Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in Brazil.
• Design and oversee the creation of diverse integration pipelines utilizing Apache Airflow, Dagster, Airbyte, AWS DMS, dbt, Kubernetes, REST APIs, SFTP, and various connectors for ongoing data ingestion.
• Employ advanced techniques for partitioning and indexing (Z-Ordering/Liquid Clustering), optimize file layouts, and implement effective read/write strategies on Amazon S3, Databricks, and dbt.
• Minimize excessive table scans (full table scans), decrease S3 request costs (GET/LIST), and enhance query performance.
• Create and execute real-time, topic- and event-driven data ingestion using Apache Kafka, AWS Kinesis, and Event Hubs.
• Lessen direct reliance on queries and bulk loads against relational databases like PostgreSQL and MySQL.
• Develop tools, agents, and GenAI/LLM-based pipelines to automate engineering operational tasks.
• Construct and maintain resilient pipelines employing dimensional modeling and Medallion architecture.
• Provide clean, aggregated, and optimized datasets for ingestion and consumption.
• Ensure adherence to compliance (LGPD/GDPR), implement sensitive data masking, maintain data lineage, enforce granular access control, and monitor data freshness and quality.
• Collect requirements from various business areas, document architectures, establish data contracts, and ensure governance and maintainability of pipelines.
• Extensive experience in data engineering within high-volume, production-critical settings.
• Proficient hands-on experience with partitioning strategies, data layout optimization, compression, Z-Ordering, and query optimization.
• Familiarity with Airflow, Dagster, Airbyte, AWS DMS, dbt, Kubernetes, REST APIs, SFTP, and custom connectors.
• Experience in event/topic-based ingestion architectures utilizing Kafka, Kinesis, Pulsar, or RabbitMQ.
• Advanced skills in Python, PySpark, and SQL optimized for large-scale distributed processing.
• Significant experience with Databricks, Delta Lake, Auto Loader, Structured Streaming, and Jobs/Workflows following Medallion Architecture.
• Proficient understanding of AWS data services: S3, IAM, Redshift, DynamoDB, Data Catalog, and AWS DMS.
• Practical knowledge of, or a strong interest in, tools/LLMs for automating development and data engineering tasks.
• Experience with automated testing, Git/GitHub, code reviews (PRs), and CI/CD pipelines.
• Solid understanding of dimensional data modeling (Star/Snowflake) and security/privacy guidelines (LGPD/GDPR).
• Nice to have: experience in building pipelines to support AI, chunking, embeddings, vector search, feature stores, and Unity Catalog.
• Nice to have: familiarity with Unity Catalog, Terraform, and Databricks Asset Bundles (DAB).
• Nice to have: experience in implementing data contracts and quality/observability frameworks such as Great Expectations, Soda, dbt tests, and Datadog.
• Nice to have: knowledge in cloud/Databricks cost management, anomaly alerting, and statistical detection of data drift.
• Nice to have: experience with idempotency, balance reconciliation, and financial reprocessing.
• Nice to have: background in fintech, payments, or healthcare, including compliance and data audits.
• Caju Card, providing greater flexibility to utilize your benefits (Meal, Food, Mobility, Health, Home Office, Culture, and Education).
• Health plan with no copayment (Unimed, Sulamerica, or Alice).
• Zenklub: online consultations with therapists and coaches to support mental health.
• Wellhub.
• Support for language learning through a partnership with Rosetta Stone.
• Recharge day - an extra day off.
• Conexa Saúde - online medical consultations.
• Childcare assistance.
• Partnership with Alura.
• Remote work — the ability to work from anywhere within Brazil.
• Work equipment provided.
• Numerous growth opportunities.
Pluribus Digital
GoMining
Get handpicked remote jobs straight to your inbox weekly.