AWS Lakehouse Data Engineer

atGuidehouseRemoteUS flagUnited StatesFull-timeData EngineerMid-levelSenior$113k – $188k/year

Posted Aug 31

This is a fully remote position, open to applicants in United States.

📋 Description

• Design, develop, and manage a cloud-native data platform that supports AI/ML, analytics, reporting, and data visualization.

• Create and execute batch and streaming ingestion processes from APIs, relational databases, file uploads, event streams, and external partners.

• Develop, test, and enhance Python and PySpark ETL/ELT pipelines for analytics-ready datasets.

• Implement incremental processing, change data capture, data contracts, schema validation, and reusable transformation frameworks.

• Enhance pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational documentation.

• Design and build a Delta Lakehouse-style platform utilizing AWS-native services.

• Construct and oversee a scalable lakehouse on Amazon S3 using Apache Iceberg and Apache Parquet.

• Implement ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities.

• Facilitate rapid interactive querying with tools like Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift.

• Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and separation of compute and storage.

• Establish standardized environments for development, testing, and production with controlled promotion processes.

• Implement governance, metadata management, lineage tracking, fine-grained access control, classification, retention, encryption, and secure data handling practices.

• Create operational data quality checks and publish measurable SLAs/SLOs.

• Facilitate AWS provisioning utilizing Infrastructure as Code and secure-by-default baselines.

• Develop and refine CI/CD processes for data pipelines and lakehouse components, encompassing testing, security checks, deployment, promotion, and rollback.

• Incorporate observability through metrics, logs, traces, alerts, dashboards, runbooks, and incident response procedures.

• Assess and enhance platform performance, scalability, reliability, security, and cost efficiency.

• Collaborate with teams across data, applications, analytics, AI/ML, security, networking, and cloud platforms.

• Maintain architecture diagrams, data models, SOPs, interface specifications, runbooks, and secure configuration baselines.

• Communicate technical findings, trade-offs, risks, and recommendations to both technical and non-technical stakeholders.


⛳️ Requirements

• Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years of equivalent practical experience in lieu of a degree.

• SIX (6) years of relevant professional experience.

• Hands-on expertise in implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services like AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.

• Strong background in developing production ETL/ELT pipelines with Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.

• Practical experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.

• Advanced SQL skills and experience with analytical queries, semantic layers, reporting tools, and data visualization workloads.

• Experience in implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.

• Knowledge of AWS security fundamentals, such as IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.

• Experience provisioning AWS resources using Infrastructure as Code and managing data platforms across multiple environments.

• Familiarity with building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.

• Ability to diagnose issues in distributed data-processing workloads and optimize their performance, reliability, and cost efficiency.

• Capability to obtain Public Trust clearance.


🏝️ Benefits

• Medical, Rx, Dental & Vision Insurance.

• Personal and Family Sick Time & Company Paid Holidays.

• Parental Leave.

• 401(k) Retirement Plan.

• Group Term Life and Travel Assistance.

• Voluntary Life and AD&D Insurance.

• Health Savings Account, Health Care & Dependent Care Flexible Spending Accounts.

• Transit and Parking Commuter Benefits.

• Short-Term & Long-Term Disability.

• Tuition Reimbursement, Personal Development, Certifications & Learning Opportunities.

• Employee Referral Program.

• Corporate Sponsored Events & Community Outreach.

• Care.com annual membership.

• Employee Assistance Program.

• Supplemental Benefits via Corestream (Critical Care, Hospital Indemnity, Accident Insurance, Legal Assistance, and ID theft protection, etc.).

• Position may be eligible for a discretionary variable incentive bonus.

• Flexible benefits package.

People also viewed

ASRC Federal1 day ago

Principal Data Architect

US flagDistrict of Columbia, +1 more stateFull-timeData Engineer
ApplyView job
Sedona Digital1 day ago

Senior Data Engineer

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
Everwest1 day ago

Lead Data Engineer

LT flagLithuania OnlyFull-timeData Engineer€4,500 – €6,500/month
ApplyView job
Distrito1 day ago

Data Engineer

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
EVT1 day ago

Senior Data Engineer

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job
Booker DiMaio1 day ago

Senior Oracle / Informatica Data Warehouse Engineer

US flagUnited States OnlyFull-timeData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers