
Master Data Developer
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in Brazil.
• Design, develop, and sustain robust ETL/ELT workflows to ingest, transform, and deliver data within a contemporary Data Lake framework.
• Create and enhance distributed data processing workflows using Python and PySpark to efficiently manage large-scale datasets.
• Establish and improve partitioning strategies for data lake storage systems (such as Delta Lake or Apache Iceberg) to optimize query performance while managing storage expenses.
• Compose, optimize, and convert intricate SQL queries that incorporate CTEs, window functions, conditional expressions, and aggregations.
• Transition and modernize data pipelines from traditional RDBMS systems to cloud-native analytical environments.
• Engage confidently with AWS-native services including Glue (Jobs, Catalog, Triggers, Workflows), Athena, Redshift, S3, Lambda, EventBridge, and other relevant data services.
• Collaborate with infrastructure and DevOps teams to provision and oversee data resources using Infrastructure as Code (IaC) tools such as CloudFormation, CDK, or Terraform.
• Monitor the health and performance of data pipelines utilizing CloudWatch and other observability tools, actively resolving issues and enhancing reliability.
• Ensure data integrity, consistency, and regulatory compliance throughout pipelines and storage layers.
• Partner with data analysts, scientists, and business stakeholders to comprehend requirements and convert them into scalable technical solutions.
• Remain updated with the latest data engineering practices, tools, and innovations in cloud-native technologies.
• Strong experience in ETL processes and data pipeline development utilizing AWS.
• High proficiency in Python as the main programming language, with proven experience in writing and optimizing PySpark code for distributed data processing.
• Comprehensive understanding of SQL, including complex queries (CTEs, window functions, aggregations, conditional expressions) and experience in translating workloads from legacy RDBMS systems.
• Practical experience with AWS Glue (Jobs, Catalog, Triggers, Workflows), Athena, and Redshift.
• Solid grasp of Data Lake architectures and partitioning strategies to enhance performance and manage costs.
• Good understanding of object-oriented programming (OOP) principles and experience with reusable code libraries.
• Proficient in using Git, Shell scripts, and Linux environments.
• Familiarity with observability, monitoring, and metric tracking methodologies.
• Advanced/Fluent proficiency in English.
• Health and dental insurance.
• Meal and food allowance.
• Childcare assistance.
• Extended paternity leave.
• Partnerships with gyms and health and wellness professionals through Wellhub (Gympass) TotalPass.
• Profit Sharing and Results Participation (PLR).
• Life insurance.
• Continuous learning platform (CI&T University).
• Discount club.
• Free online platform focused on physical, mental, and overall well-being.
• Pregnancy and responsible parenting course.
• Collaborations with online learning platforms.
• Language learning platform.
• And many more!
Social Discovery Group
NatWest Group
Towa Software
Clinical Outcomes Solutions
Get handpicked remote jobs straight to your inbox weekly.