
Data Platform Engineer
Posted 11 hours ago

Posted 11 hours ago
This is a fully remote position, open to applicants in Brazil.
β’ Transform Azure Data Factory pipelines into Databricks workflows on AWS utilizing reusable templates.
β’ Rehost Databricks workspaces on AWS and transition ADLS Gen2 storage to S3.
β’ Rewrite ADF Web Activities as Lambda functions or Step Functions tasks.
β’ Substitute ADF-specific scaling with native mechanisms from Databricks.
β’ Construct and optimize PySpark transformations for production-level data volumes.
β’ Replace Azure Synapse Serverless with Databricks SQL Warehouse.
β’ Reconcile the migrated data with source systems.
β’ Migrate 2,094 ADF pipelines, convert 9,859 pipeline activities, transfer 471 Spark dataflows to Databricks on AWS, and rehost four Databricks workspaces across Development, QA, Pre-production, and Production environments.
β’ Proven production experience with Databricks, encompassing workspaces, jobs, and workflows.
β’ Strong background in Spark and PySpark, particularly with real data volumes and performance tuning.
β’ Proficiency in production-grade Python.
β’ Experience in building or migrating Azure Data Factory pipelines, along with a solid understanding of the ADF activity model.
β’ Familiarity with AWS data services: S3, Glue, Athena, Lambda, and Step Functions.
β’ Advanced SQL skills, including the ability to read and interpret stored procedures.
β’ Professional proficiency in written and spoken English.
β’ Preferred: Experience with Delta Lake, Unity Catalog, Azure Synapse, Terraform, Airflow/MWAA, dbt, Kafka, Databricks certification, data modeling, Great Expectations, SAS/analytics platform integration, and CRM data.
β’ 100% remote work.
β’ Full-time engagement.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.