Data Engineer – Mid Level

Posted Sep 10

This is a fully remote position, open to applicants in India.

📋 Description

• Design, develop, and enhance batch and streaming data ingestion pipelines across AWS, Azure, and GCP.

• Create pipelines utilizing Databricks Workflows, Apache Spark/PySpark, SQL, Delta Live Tables, and Databricks Lakeflow components.

• Execute Bronze–Silver–Gold medallion architecture, CDC, SCD Type 1 and Type 2, schema evolution, validation, reconciliation, and data-quality rules.

• Set up and manage Delta Lake storage frameworks, tables, schemas, partitions, and optimization processes.

• Implement OPTIMIZE, Z-ORDER, VACUUM, partitioning, and file management strategies.

• Support Unity Catalog metadata, including cataloging, lineage, governance, and integration with Microsoft Purview.

• Integrate Amazon S3, Azure Storage, and Google Cloud Storage with Databricks.

• Establish data-quality checks, profiling, validation, monitoring, RBAC policies, security controls, and data classification tags.

• Construct, schedule, oversee, and maintain production workflows using Databricks Workflows, Delta Live Tables, Azure Data Factory, and other approved tools.

• Contribute to CI/CD pipelines, automated testing, deployment, environment management, and DEV–QA–PROD promotion.

• Monitor production pipelines, troubleshoot failures, investigate root causes, support recovery, and optimize performance.

• Collaborate with the Senior Data Architect, Data Scientists, ML Engineers, Analysts, Product teams, and engineering stakeholders.

• Engage in architecture reviews, technical design discussions, coding reviews, and engineering standards meetings.

• Document pipelines, data flows, data dictionaries, transformation logic, data-quality rules, test cases, job schedules, and operational procedures.

• Convert architecture designs and technical standards into reliable, production-ready pipelines and platform capabilities.


⛳️ Requirements

• 3–5 years of experience in Data Engineering, ETL development, or cloud data platform engineering.

• Practical experience with Databricks, Apache Spark, PySpark, or other distributed data processing technologies.

• Strong expertise in SQL, encompassing structured data transformation, joins, aggregations, and performance-aware query development.

• Familiarity with at least one major cloud platform; Microsoft Azure preferred, with AWS and/or GCP also advantageous.

• Knowledge of data modeling, data quality, schema evolution, data validation, pipeline monitoring, and troubleshooting.

• Basic understanding of RBAC, encryption, credential and secret management, and secure access to cloud and data platform resources.

• Working knowledge of Delta Lake, medallion architecture, and modern lakehouse best practices is preferred.

• Experience with metadata, cataloging, and governance platforms such as Unity Catalog, Microsoft Purview, or AWS Glue Data Catalog is preferred.

• Familiarity with workflow orchestration and scheduling technologies such as Azure Data Factory, Databricks Workflows, Apache Airflow, or similar frameworks is preferred.

• Experience with Git-based development, CI/CD, and DevOps practices is preferred.

• Knowledge or experience in geospatial/GIS data, BI semantic layers, particularly Power BI, or data preparation for AI/ML workloads is preferred.

• Relevant cloud or Databricks certifications are preferred.

• Understanding of asset integrity management concepts is a plus.

• Experience with oil & gas, utility, infrastructure, or pipeline asset data is a plus.

• Familiarity with regulatory, compliance, and audit reporting requirements is a plus.


🏝️ Benefits

• Competitive compensation package based on experience and qualifications.

• Medical, Dental, and Vision Insurance.

• 401(k) Plan with Company Match.

• Generous Paid Time Off (PTO).

• Company-Paid Holidays.

• Flexible Work Options / work-from-home opportunities, depending on role and business needs.

• On-Call Compensation for eligible on-call shifts.

People also viewed

Data Elephant11 hours ago

Software Engineer – Data Products, AI Applications

CA flagCanada OnlyFreelanceData Engineer
ApplyView job
ICF19 hours ago

Senior Data Engineer, Scala

US flagVirginia OnlyFull-timeData Engineer$98.6k – $167.6k/year
ApplyView job
General Dynamics Information Technology1 day ago

Data Architect Principal, Analytics Pipeline

US flagUnited States OnlyFull-timeData Engineer$136k – $184k/year
ApplyView job
Logic20/20, Inc.1 day ago

Lead Geospatial Data Engineer

US flagWashington OnlyFull-timeData Engineer$173.1k – $179.9k/year
ApplyView job
Rox Partner1 day ago

Senior Data Engineer

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job
EVT1 day ago

Data Engineer – GCP

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers