
Senior Data Engineer
Posted Aug 6

Posted Aug 6
This is a fully remote position, open to applicants in Brazil.
• Design, develop, and maintain comprehensive batch pipelines (ingestion, transformation, modeling, and serving) that handle terabytes of data.
• Create and optimize PySpark jobs, focusing on aspects such as partitioning, shuffle, skew, memory utilization, and execution costs.
• Develop, maintain, and enhance Directed Acyclic Graphs (DAGs) in Apache Airflow, ensuring idempotency, effective failure handling, retries, and compliance with SLAs.
• Model and implement transformations using dbt, ensuring thorough testing, documentation, lineage tracking, and adherence to good versioning practices.
• Advance data layers (raw, curated, analytics) while upholding consistent quality standards, contracts, and SLAs.
• Investigate and address incidents in production pipelines, conducting root cause analysis and recommending structural enhancements.
• Implement and sustain data quality, observability, and monitoring frameworks.
• Collaborate with analytics and business teams to comprehend requirements, propose suitable data models, and guarantee the reliability of the delivered data.
• Participate in architectural decisions, code reviews, and the sharing of best practices within the team.
• Located in São Paulo (city).
• Bachelor's degree in Computer Science, Engineering, Mathematics, or a related discipline.
• Over 5 years of strong experience in data engineering, particularly in large-scale environments.
• Proficient in Python with a focus on good coding practices, testing, and modularization.
• Extensive production experience with PySpark, including job tuning and troubleshooting at scale.
• Advanced SQL skills with expertise in window functions, Common Table Expressions (CTEs), query optimization, and analytical modeling.
• Hands-on experience with dbt on medium to large projects (incremental models, tests, macros, exposures).
• Solid production experience with Apache Airflow, including the development of complex DAGs, custom operators, sensors, dependency management, and execution troubleshooting.
• Proven experience with AWS data services (S3, Glue, EMR/EMR Serverless, Athena, IAM, Lambda, among others).
• Familiarity with open table formats (Iceberg, Delta, or Hudi) and their implications for performance and cost.
• Capability to discuss and justify architectural trade-offs (cost, latency, complexity, maintainability).
• Practical experience in applying Data Security and Compliance policies (LGPD/GDPR).
• Experience with Apache Iceberg in a production setting.
• Knowledge of Infrastructure as Code (Terraform, CDK).
• CI/CD practices applied to data projects (automated tests, DAG deployments, dbt in pipelines).
• Understanding of data contracts, data cataloging, and governance.
• Prior experience in streaming environments (Kinesis, Kafka, Flink)—though this role primarily focuses on batch processing.
• Familiarity with DuckDB or similar tools for efficient analytical processing.
• Contributions to open source projects or published technical materials.
• Competitive salary and performance-based bonuses.
• Flexible work hours and remote work options.
• Opportunities for professional development and continuous learning.
• Comprehensive health and wellness benefits.
• Collaborative and innovative work environment.
Pluribus Digital
GoMining
Get handpicked remote jobs straight to your inbox weekly.