Senior Data Engineer

Posted 1 day ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Create, construct, and uphold scalable data pipelines within a Databricks E2 environment hosted on AWS.

• Gather and analyze data from relational databases, APIs, external data sources, and real-time streaming platforms.

• Design and execute pipelines to cleanse, transform, and aggregate data for reporting and analytical purposes.

• Develop and manage the Databricks medallion architecture utilizing Bronze, Silver, and Gold layers.

• Produce source-to-target mapping documentation and perform unit testing.

• Write intricate SQL aggregations and examine data for anomalies, quality concerns, and discrepancies.

• Implement real-time and near-real-time data ingestion using AWS DMS and other AWS-native services.

• Create data-processing logic using Python and/or R with Spark, PySpark, and Pandas.

• Oversee source code, version control, and deployments utilizing GitLab and CI/CD pipelines.

• Work collaboratively with cross-functional teams in an Agile project environment.

• Ensure compliance with FISMA High and multi-tenant security, governance, and compliance standards.

• Utilize AI automation tools for the development, testing, validation, and delivery of pipelines.


⛳️ Requirements

• Demonstrated, hands-on experience as a Data Engineer or Databricks Developer creating production-grade data pipelines.

• In-depth knowledge of Databricks on AWS, including E2 architecture, cluster setup, job orchestration, and workspace management.

• Practical experience with Apache Spark, PySpark, and Pandas for large-scale, distributed data processing tasks.

• Strong programming abilities in Python; familiarity with R is advantageous.

• Experience with Databricks Auto Loader for efficient, incremental file ingestion.

• Hands-on experience with AWS Database Migration Service (DMS) for change data capture and real-time data replication.

• Extensive experience with Delta Lake and Delta tables, including schema evolution, time travel, OPTIMIZE, Z-ORDER, and VACUUM.

• Familiarity with Amazon RDS and other relational database sources utilized in data extraction and integration processes.

• Advanced SQL capabilities, including complex joins, window functions, aggregations, and performance optimization.

• Strong comprehension of medallion architecture and contemporary data lakehouse design principles.

• Experience with GitLab and CI/CD pipelines for automated testing, building, and deployment of data engineering code.

• Background in Agile/Scrum project environments.


🏝️ Benefits

• Competitive salary.

• Medical, dental, vision, life, and disability coverage.

• Optional critical illness, hospital, and accident insurance.

• Health savings and flexible spending accounts.

• Retirement 401K plan.

• Paid leave programs.

• Flexible work-life balance.

• Opportunities for professional growth.

• Performance and recognition initiatives.

• Collaborative workplace culture.

• Financial and non-financial benefits for Seneca Nation members.

People also viewed

CI&T1 day ago

Senior Data Architect

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
Tonic31 day ago

Data Engineer – Microsoft Fabric, Data Pipelines

AR flagArgentina, +1 more countryPart-timeData Engineer
ApplyView job
CI&T1 day ago

Senior Data Architect

CO flagColombia OnlyFull-timeData Engineer
ApplyView job
Latitude IT Solutions | SDVOSB1 day ago

Data Engineer

US flagUnited States OnlyFull-timeData Engineer$142k – $158k/year
ApplyView job
CipherHealth1 day ago

Data Engineer

US flagUnited States OnlyFull-timeData Engineer
ApplyView job
Civic Marketplace1 day ago

Data Engineer

GB flagUnited Kingdom, +1 more countryFull-timeData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers