Remotery

Principal Software Engineer – Data Platform, Iceberg/Trino

Posted Aug 14

This is a fully remote position, open to applicants in United States.

📋 Description

• Take ownership of the lakehouse reference architecture, which includes the design of Iceberg tables, Trino cluster topology, catalog services, Spark transformation compute, and object storage layout.

• Develop on-premise alternatives for cloud-managed warehouse features, such as change-data-capture streams, scheduled tasks, and write-back pathways to operational stores.

• Conduct proof-of-concept validations for the catalog and query engine at anticipated data volumes.

• Establish evidence-based criteria for making decisions between VM-based and Kubernetes-native operator placements.

• Set organization-wide standards for table configurations, partitioning, file sizing, and Iceberg maintenance, which encompasses compaction, snapshot expiry, and orphan-file management.

• Lead the SQL dialect strategy for migrating existing warehouse workloads to Trino and Spark SQL.

• Guide senior engineers across various data workstreams and evaluate design proposals.

• Collaborate with platform engineering on storage sizing, resource isolation, and lakehouse capacity planning.


⛳️ Requirements

• B.E., B.Tech., or M.Sc. degree in Computer Science or a related technical discipline.

• Over 12 years of industry experience in constructing and managing large-scale data platforms or distributed systems.

• Extensive hands-on expertise in distributed SQL engines, including Trino/Presto or Spark SQL internals, query planning, and performance optimization.

• Production-level experience with Apache Iceberg, or Delta Lake/Hudi, along with a willingness to gain in-depth knowledge of Iceberg.

• Understanding of Iceberg table specifications, including merge-on-read versus copy-on-write, and large-scale table maintenance.

• Familiarity with Iceberg catalog services, encompassing REST catalogs like Polaris or Nessie, or Hive Metastore.

• Knowledge of S3-compatible object storage solutions.

• Solid comprehension of cloud warehouse internals, including Snowflake, BigQuery, or Redshift.

• Professional software development experience in Java and/or Python.

• Experience in delivering data platforms within on-premise, regulated, or air-gapped environments is a significant advantage.

• Experience with healthcare data is a plus.


🏝️ Benefits

• 20 days of fixed paid time off annually, in addition to company holidays.

• Generous parental leave policy.

• Monetary incentives and company-wide recognition for contributions and dedication.

• Medical, dental, and vision insurance coverage.

• 100% company-funded short- and long-term disability insurance.

• 100% company-funded basic life insurance.

• Discounted legal assistance.

• Pet insurance options available.

People also viewed

Railroad196 hours ago

Senior Data Engineer – GCP, Python, Iceberg, Delta Lake, Kafka, Snowflake, Databricks

US flagUnited States OnlyFull-timeData Engineer$120k – $180k/year
ApplyView job
Livefront6 hours ago

Data Engineer

PE flagPeru OnlyFull-timeData Engineer
ApplyView job
GFT Technologies6 hours ago

Data Engineer, Mid-level

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
VIDA7 hours ago

Geospatial Data Engineer – Customer & AI Solutions

DE flagGermany OnlyFull-timeData Engineer
ApplyView job
albo7 hours ago

Data Engineer

MX flagMexico OnlyFull-timeData Engineer
ApplyView job
Leega7 hours ago

Engenheiro de Dados Pleno – AWS

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers