
AWS Lakehouse Data Engineer
Posted Aug 31

Posted Aug 31
This is a fully remote position, open to applicants in United States.
• Design, develop, and manage a cloud-native data platform that supports AI/ML, analytics, reporting, and data visualization.
• Create and execute batch and streaming ingestion processes from APIs, relational databases, file uploads, event streams, and external partners.
• Develop, test, and enhance Python and PySpark ETL/ELT pipelines for analytics-ready datasets.
• Implement incremental processing, change data capture, data contracts, schema validation, and reusable transformation frameworks.
• Enhance pipeline reliability through automated testing, orchestration, monitoring, retry handling, and operational documentation.
• Design and build a Delta Lakehouse-style platform utilizing AWS-native services.
• Construct and oversee a scalable lakehouse on Amazon S3 using Apache Iceberg and Apache Parquet.
• Implement ACID transactions, schema evolution, partition evolution, snapshot isolation, time travel, and rollback capabilities.
• Facilitate rapid interactive querying with tools like Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift.
• Optimize performance and cost through partitioning, compaction, file sizing, statistics, caching, lifecycle policies, and separation of compute and storage.
• Establish standardized environments for development, testing, and production with controlled promotion processes.
• Implement governance, metadata management, lineage tracking, fine-grained access control, classification, retention, encryption, and secure data handling practices.
• Create operational data quality checks and publish measurable SLAs/SLOs.
• Facilitate AWS provisioning utilizing Infrastructure as Code and secure-by-default baselines.
• Develop and refine CI/CD processes for data pipelines and lakehouse components, encompassing testing, security checks, deployment, promotion, and rollback.
• Incorporate observability through metrics, logs, traces, alerts, dashboards, runbooks, and incident response procedures.
• Assess and enhance platform performance, scalability, reliability, security, and cost efficiency.
• Collaborate with teams across data, applications, analytics, AI/ML, security, networking, and cloud platforms.
• Maintain architecture diagrams, data models, SOPs, interface specifications, runbooks, and secure configuration baselines.
• Communicate technical findings, trade-offs, risks, and recommendations to both technical and non-technical stakeholders.
• Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or a related field, or FOUR (4) years of equivalent practical experience in lieu of a degree.
• SIX (6) years of relevant professional experience.
• Hands-on expertise in implementing AWS-native data lake or lakehouse architectures using Amazon S3 and services like AWS Glue, Amazon Athena, Amazon EMR, AWS Lake Formation, and Amazon Redshift.
• Strong background in developing production ETL/ELT pipelines with Python and PySpark, including data modeling, transformation, testing, performance tuning, and error handling.
• Practical experience with Apache Iceberg, including ACID transactions, snapshots, schema and partition evolution, time travel, table maintenance, and query optimization.
• Advanced SQL skills and experience with analytical queries, semantic layers, reporting tools, and data visualization workloads.
• Experience in implementing metadata management and governance capabilities, including cataloging, lineage, ownership, classification, policy enforcement, and fine-grained access controls.
• Knowledge of AWS security fundamentals, such as IAM and least privilege, KMS encryption, secrets management, network security, logging, and secure SDLC practices.
• Experience provisioning AWS resources using Infrastructure as Code and managing data platforms across multiple environments.
• Familiarity with building or operating CI/CD pipelines for data workflows, including testing, packaging, deployment automation, environment promotion, and rollback.
• Ability to diagnose issues in distributed data-processing workloads and optimize their performance, reliability, and cost efficiency.
• Capability to obtain Public Trust clearance.
• Medical, Rx, Dental & Vision Insurance.
• Personal and Family Sick Time & Company Paid Holidays.
• Parental Leave.
• 401(k) Retirement Plan.
• Group Term Life and Travel Assistance.
• Voluntary Life and AD&D Insurance.
• Health Savings Account, Health Care & Dependent Care Flexible Spending Accounts.
• Transit and Parking Commuter Benefits.
• Short-Term & Long-Term Disability.
• Tuition Reimbursement, Personal Development, Certifications & Learning Opportunities.
• Employee Referral Program.
• Corporate Sponsored Events & Community Outreach.
• Care.com annual membership.
• Employee Assistance Program.
• Supplemental Benefits via Corestream (Critical Care, Hospital Indemnity, Accident Insurance, Legal Assistance, and ID theft protection, etc.).
• Position may be eligible for a discretionary variable incentive bonus.
• Flexible benefits package.
ASRC Federal
Get handpicked remote jobs straight to your inbox weekly.