
Senior Data Engineer
Posted Jul 23

Posted Jul 23
This is a fully remote position, open to applicants in Pakistan.
• Take charge of deploying, scaling, and orchestrating the core data stack utilizing AWS Managed Workflows for Apache Airflow (MWAA).
• Craft and enhance the overall architecture of the data platform, carefully considering scalability, reliability, cost, and maintainability.
• Configure and optimize raw data storage within S3, overseeing Redshift external schemas (Spectrum) and the AWS Glue Data Catalog.
• Establish effective partitioning strategies, utilize columnar file formats (Parquet), and optimize storage arrangements to enhance query performance, scalability, and cost efficiency.
• Develop scalable analytical data models using dimensional modeling and lakehouse best practices to cater to reporting, analytics, and downstream users.
• Oversee and enhance dbt to guarantee efficient transformation, testing, and materialization of raw data within Amazon Redshift.
• Incorporate observability into the data platform via monitoring, logging, alerting, and operational dashboards.
• Assist in designing and implementing the shift toward event-based ingestion into S3.
• Set up and maintain CI/CD deployment pipelines for dbt projects, Airflow DAGs, and infrastructure.
• Ensure optimal performance, access control, and uptime for BI tools linked to Redshift.
• In-depth hands-on experience with deploying and managing AWS data services, particularly MWAA (Airflow), Redshift / Redshift Spectrum, S3, IAM, and Glue.
• Advanced expertise in dbt (structuring dbt projects, configuring sources, writing custom macros, and optimizing incremental models).
• Strong practical experience in deploying AWS data platform components using Terraform or AWS CloudFormation.
• Expert-level SQL capabilities, with a thorough understanding of Redshift distribution/sort keys and query optimization across external schemas.
• Proficient in Python for developing Airflow DAGs, automation, integrations, and data engineering tools.
• Experience in designing efficient data lakes using Parquet, partitioning strategies, metadata catalogs, and external table technologies like Redshift Spectrum.
• Familiarity with open table formats such as Apache Iceberg, Delta Lake, or Apache Hudi, including knowledge of ACID transactions, schema evolution, time travel, and metadata management.
• Experience with Apache Kafka or AWS MSK for event-driven data streaming and real-time ingestion into S3 (Bonus).
• Experience in designing resilient streaming pipelines with suitable delivery guarantees, reconciliation, backfill strategies, and schema contract management (Nice-to-Have).
• Health insurance
• Flexible work arrangements
• Professional development opportunities
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.