
Senior Data Engineer
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Create, construct, and uphold scalable data pipelines within a Databricks E2 environment hosted on AWS.
• Gather and analyze data from relational databases, APIs, external data sources, and real-time streaming platforms.
• Design and execute pipelines to cleanse, transform, and aggregate data for reporting and analytical purposes.
• Develop and manage the Databricks medallion architecture utilizing Bronze, Silver, and Gold layers.
• Produce source-to-target mapping documentation and perform unit testing.
• Write intricate SQL aggregations and examine data for anomalies, quality concerns, and discrepancies.
• Implement real-time and near-real-time data ingestion using AWS DMS and other AWS-native services.
• Create data-processing logic using Python and/or R with Spark, PySpark, and Pandas.
• Oversee source code, version control, and deployments utilizing GitLab and CI/CD pipelines.
• Work collaboratively with cross-functional teams in an Agile project environment.
• Ensure compliance with FISMA High and multi-tenant security, governance, and compliance standards.
• Utilize AI automation tools for the development, testing, validation, and delivery of pipelines.
• Demonstrated, hands-on experience as a Data Engineer or Databricks Developer creating production-grade data pipelines.
• In-depth knowledge of Databricks on AWS, including E2 architecture, cluster setup, job orchestration, and workspace management.
• Practical experience with Apache Spark, PySpark, and Pandas for large-scale, distributed data processing tasks.
• Strong programming abilities in Python; familiarity with R is advantageous.
• Experience with Databricks Auto Loader for efficient, incremental file ingestion.
• Hands-on experience with AWS Database Migration Service (DMS) for change data capture and real-time data replication.
• Extensive experience with Delta Lake and Delta tables, including schema evolution, time travel, OPTIMIZE, Z-ORDER, and VACUUM.
• Familiarity with Amazon RDS and other relational database sources utilized in data extraction and integration processes.
• Advanced SQL capabilities, including complex joins, window functions, aggregations, and performance optimization.
• Strong comprehension of medallion architecture and contemporary data lakehouse design principles.
• Experience with GitLab and CI/CD pipelines for automated testing, building, and deployment of data engineering code.
• Background in Agile/Scrum project environments.
• Competitive salary.
• Medical, dental, vision, life, and disability coverage.
• Optional critical illness, hospital, and accident insurance.
• Health savings and flexible spending accounts.
• Retirement 401K plan.
• Paid leave programs.
• Flexible work-life balance.
• Opportunities for professional growth.
• Performance and recognition initiatives.
• Collaborative workplace culture.
• Financial and non-financial benefits for Seneca Nation members.
Tonic3
Latitude IT Solutions | SDVOSB
Get handpicked remote jobs straight to your inbox weekly.