
Mid-Level Data Engineer – GCP, DBT
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in Brazil.
• Evaluate the architecture and requirements of data warehouses
• Map data, transformations, and processes across Google Cloud Platform (GCP) services, such as Cloud Storage, BigQuery, and Dataproc
• Develop data migration strategies, including full load, incremental load, and Change Data Capture (CDC)
• Create a comprehensive GCP data architecture plan
• Design table schemas in BigQuery with a focus on performance, cost-effectiveness, and scalability
• Establish partitioning and clustering strategies for BigQuery
• Model Bronze, Silver, and Gold data zones within Cloud Storage
• Construct transformation routines utilizing Dataproc/Spark or Dataflow to populate data into BigQuery
• Convert business logic and existing transformations to fit within GCP
• Execute data validation and quality assurance processes
• Enhance BigQuery queries, optimize Spark jobs in Dataproc, and efficiently utilize GCP resources
• Implement data security measures for both in transit and at rest
• Define and enforce Identity and Access Management (IAM) policies
• Ensure adherence to data governance policies
• Diagnose performance and functionality issues within pipelines and GCP resources
• Document architecture, pipelines, data models, and operational procedures
• Engage with team members, stakeholders, and other departments
• Maintain clear communication regarding architecture, software components, development progress, and quality standards
• Utilize Agile methodologies and manage tasks using Jira
• At least 3 years of proven experience with DBT
• Strong understanding of models such as staging, intermediate, and marts
• Familiarity with ref() and source() functions
• Knowledge of macros, specifically Jinja
• Understanding of seeds and snapshots
• Proficiency in not null, unique, and custom tests
• Experience with layered organization: Staging → Transform → Mart
• Extensive knowledge of BigQuery, including data modeling, query optimization, partitioning, clustering, streaming and batch loads, security, and governance
• Experience with Cloud Storage, including buckets, storage classes, lifecycle policies, IAM, and security measures
• Capability to provision, configure, and manage Spark/Hadoop clusters in Dataproc
• Familiarity with job optimization and integration with additional GCP services
• Knowledge of Dataflow, Composer, and DBT for orchestrating and processing data
• Understanding of Cloud IAM and granular access control
• Knowledge of Virtual Private Clouds (VPCs), networking, subnets, firewall rules, and cloud security
• Proficiency in Python and PySpark
• Advanced SQL skills
• Shell scripting expertise
• Experience with Git/GitHub/Bitbucket
• Familiarity with Agile methodologies, ceremonies, and proficiency with Jira
• Porto Seguro health insurance, with options to include a spouse and children
• Porto Seguro dental insurance for employees and dependents
• Profit Sharing and Results (PLR)
• Childcare assistance
• Alelo food and meal vouchers
• Home office allowance
• Partnerships with educational institutions, providing discounts and incentives for courses and degree programs
• Certification incentives for cloud certifications (GCP, Azure, AWS, and others)
• Livelo points
• TotalPass
• Mindself, offering incentives for meditation and mindfulness practices
• Continuous professional development opportunities
Truv
Stellar Health
LIBBS FARMACÊUTICA LTDA
Leega
Get handpicked remote jobs straight to your inbox weekly.