
Principal Software Engineer – Data Platform, Iceberg/Trino
Posted Aug 14

Posted Aug 14
This is a fully remote position, open to applicants in United States.
• Take ownership of the lakehouse reference architecture, which includes the design of Iceberg tables, Trino cluster topology, catalog services, Spark transformation compute, and object storage layout.
• Develop on-premise alternatives for cloud-managed warehouse features, such as change-data-capture streams, scheduled tasks, and write-back pathways to operational stores.
• Conduct proof-of-concept validations for the catalog and query engine at anticipated data volumes.
• Establish evidence-based criteria for making decisions between VM-based and Kubernetes-native operator placements.
• Set organization-wide standards for table configurations, partitioning, file sizing, and Iceberg maintenance, which encompasses compaction, snapshot expiry, and orphan-file management.
• Lead the SQL dialect strategy for migrating existing warehouse workloads to Trino and Spark SQL.
• Guide senior engineers across various data workstreams and evaluate design proposals.
• Collaborate with platform engineering on storage sizing, resource isolation, and lakehouse capacity planning.
• B.E., B.Tech., or M.Sc. degree in Computer Science or a related technical discipline.
• Over 12 years of industry experience in constructing and managing large-scale data platforms or distributed systems.
• Extensive hands-on expertise in distributed SQL engines, including Trino/Presto or Spark SQL internals, query planning, and performance optimization.
• Production-level experience with Apache Iceberg, or Delta Lake/Hudi, along with a willingness to gain in-depth knowledge of Iceberg.
• Understanding of Iceberg table specifications, including merge-on-read versus copy-on-write, and large-scale table maintenance.
• Familiarity with Iceberg catalog services, encompassing REST catalogs like Polaris or Nessie, or Hive Metastore.
• Knowledge of S3-compatible object storage solutions.
• Solid comprehension of cloud warehouse internals, including Snowflake, BigQuery, or Redshift.
• Professional software development experience in Java and/or Python.
• Experience in delivering data platforms within on-premise, regulated, or air-gapped environments is a significant advantage.
• Experience with healthcare data is a plus.
• 20 days of fixed paid time off annually, in addition to company holidays.
• Generous parental leave policy.
• Monetary incentives and company-wide recognition for contributions and dedication.
• Medical, dental, and vision insurance coverage.
• 100% company-funded short- and long-term disability insurance.
• 100% company-funded basic life insurance.
• Discounted legal assistance.
• Pet insurance options available.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.