
Senior Data Engineer, Data Lakehouse Infrastructure
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in North America.
• Design and expand a high-performance data lakehouse on GCP, utilizing technologies such as StarRocks, Apache Iceberg, GCS, BigQuery, Dataproc, and Kafka.
• Create, construct, and enhance distributed query engines like Trino, Spark, or Snowflake to accommodate intricate analytical tasks.
• Execute metadata management using open table formats like Iceberg and establish data discovery frameworks for governance and observability through Iceberg-compatible catalogs.
• Build and manage robust ETL/ELT pipelines with Apache Airflow, Spark, and GCP-native tools (e.g., Dataflow, Composer).
• Work collaboratively across various teams, partnering with data scientists, backend engineers, and product managers to design and execute solutions.
• Over 5 years of experience in data or software engineering, specializing in distributed data systems and cloud-native architectures.
• Demonstrated expertise in constructing and scaling data platforms on GCP, encompassing storage, compute, orchestration, and monitoring.
• Strong proficiency in one or more query engines such as Trino, Presto, Spark, or Snowflake.
• Familiarity with modern table formats like Apache Hudi, Iceberg, or Delta Lake.
• Outstanding programming capabilities in Python, along with proficiency in SQL or SparkSQL.
• Practical experience in orchestrating workflows using Airflow and building streaming/batch pipelines using GCP-native services.
• Participate in TRM’s equity plan
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.