
Senior Data Engineer
Posted Sep 11

Posted Sep 11
This is a fully remote position, open to applicants in United States.
• Design, create, and sustain scalable data pipelines on AWS utilizing S3, DMS, Glue, Lambda, Step Functions/MWAA, and Redshift.
• Develop batch and near-real-time ETL/ELT workflows to ingest, cleanse, transform, and load data from databases, legacy systems, and event streams employing Python and PySpark.
• Create incremental/CDC mechanisms featuring restartability, idempotency, duplicate handling, and recovery.
• Execute automated controls for completeness, accuracy, reconciliation, schema alterations, and lineage tracking.
• Enhance Glue/Spark, Athena, Redshift, and S3 workloads via partitioning, columnar formats, query optimization, and appropriate storage/compute design.
• Construct near-real-time and event-driven pipelines utilizing Kinesis/Kafka, including ordering, retries, idempotency, and failure recovery mechanisms.
• Implement AWS data security protocols, least-privilege access, data classification, and governance measures.
• Monitor pipelines using CloudWatch, troubleshoot failures, and address production data incidents.
• Uphold Git/version control, code review, automated testing, and CI/CD methodologies.
• Collaborate with product owners, architects, reporting teams, and business stakeholders to convert requirements into scalable data solutions.
• Document pipelines and operational processes.
• Bachelor's degree in a computer-related discipline from an accredited institution.
• A minimum of five (5) years of experience in data engineering, specifically in constructing scalable and distributed ETL data pipelines within enterprise settings.
• Proficient in building and managing scalable AWS-based data platforms and pipelines using Lambda, Glue, Athena, S3, Redshift, DMS, MWAA (Airflow), and Step Functions.
• Experience in supporting batch, CDC, and near real-time data processing.
• Advanced skills in Python, SQL, and PySpark.
• Practical experience in creating reusable ETL/ELT frameworks, data warehouses, data marts, and integrations across databases, APIs, event streams, and analytics environments.
• Familiarity with implementing data quality, governance, and optimization best practices.
• Experience with automated validation frameworks, Lake Formation, Glue Data Catalog, performance tuning, and cost optimization across AWS data services.
• Strong communication abilities to convey complex data concepts to business stakeholders.
• Experience in healthcare, life sciences, or other highly regulated sectors with compliance requirements such as HIPAA, GDPR, or FDA is preferred.
• Knowledge of metadata management, data lineage, data observability, master data management, or enterprise data catalog solutions.
• Understanding of data modeling, analytical data models, schema design, schema evolution, and data structures optimized for reporting and analytics.
• Familiarity with data lake and data warehouse architecture, including data partitioning and columnar storage formats such as Parquet.
• Relevant AWS certification, such as AWS Certified Data Engineer – Associate, or an equivalent cloud or data engineering certification.
• Flexible work hours in a dynamic and collaborative environment.
• Remote work options available.
• A reliable internet connection is required for remote work.
• Opportunity to travel as necessary for company meetings.
Rogon Technologies GmbH
SSC HR Solutions
SSC HR Solutions
Data Elephant
Get handpicked remote jobs straight to your inbox weekly.