Data Engineer – Mid Level

Posted Sep 10

This is a fully remote position, open to applicants in United States.

📋 Description

• Develop and sustain data ingestion pipelines for both structured and semi-structured sources, such as GIS, inline inspection data, SCADA, maintenance systems, and enterprise record systems.

• Execute batch and streaming data ingestion utilizing Databricks Workflows, Spark, PySpark, SQL, and declarative pipeline tools.

• Employ Bronze, Silver, and Gold medallion architecture patterns for data transformation, standardization, and enrichment.

• Implement change data capture, manage slowly changing dimensions, support schema evolution, and set data-validation rules.

• Normalize third-party and public data feeds, including historical weather data, soil characteristics, satellite-derived information, and one-call ticket data.

• Construct AI-assisted pipelines for the normalization of units, schemas, and semantics across disparate customer data.

• Develop automated data-quality repair workflows with provenance for synthesized values.

• Productionize document-extraction pipelines in collaboration with data scientists.

• Establish human-in-the-loop review and exception workflows for low-confidence extractions.

• Configure and manage Delta Lake tables, including partitioning strategies and optimization routines.

• Implement metadata, lineage, and cataloging standards through Unity Catalog.

• Build and maintain connectors to customer systems of record with customizable refresh schedules.

• Support geospatial data processing, including spatial joins and alignment with pipeline centerline geometry.

• Implement data-quality tests, profiling, drift monitoring, access controls, security protocols, classification tags, and regulatory traceability.

• Create, schedule, and oversee workflows; manage alerting and handling of pipeline failures.

• Contribute to CI/CD processes for pipeline code, encompassing version control, automated testing, and environment promotion.

• Resolve production incidents, recover failed pipeline executions, and optimize performance and infrastructure expenses.

• Translate Data Architect designs into practical implementations and identify design gaps or ambiguities.

• Engage in architecture, design, and code reviews.

• Document pipelines, transformation logic, data dictionaries, job schedules, operational procedures, and runbooks.

• Deliver dependable Bronze, Silver, and Gold data flows, reduce the manual onboarding effort for customer data, maintain high data-quality pass rates, comply with governance standards, minimize incidents, and collaborate effectively with architects, data scientists, application engineers, and cross-functional partners.


⛳️ Requirements

• 3–5 years of experience in data engineering, ETL development, or cloud data platform engineering.

• Practical experience with Databricks, Spark, PySpark, or similar distributed data-processing technologies.

• Proficient SQL skills and experience in structured data transformation.

• Familiarity with at least one major cloud platform; Azure experience is preferred.

• Knowledge of data modeling, data-quality practices, schema evolution, and pipeline troubleshooting.

• Experience with workflow orchestration and scheduling frameworks.

• Understanding of essential data-security practices, including access control, encryption, and credential management.

• Experience with Git-based development and comfort collaborating within a code-reviewed engineering team.

• Preferred experience with Delta Lake, medallion architecture, and lakehouse engineering best practices.

• Experience with Unity Catalog, Microsoft Purview, or similar metadata and data-lineage tools is preferred.

• Experience building pipelines for the ingestion of unstructured or semi-structured documents is preferred.

• Familiarity with geospatial data processing and common GIS data formats is preferred.

• CI/CD and DevOps experience for data workloads, including infrastructure as code (IaC), is preferred.

• Experience in preparing and transforming data specifically for machine learning or probabilistic model consumption is preferred.

• Cloud or Databricks certifications are preferred.

• Experience using AI-assisted coding tools such as Cursor or GitHub Copilot and/or agentic coding tools like Claude Code within a professional development workflow is preferred.

• Experience integrating oil and gas or utility asset data, including pipelines, facilities, and GIS assets, is a nice-to-have.

• Understanding of asset integrity concepts, including inspection data, risk scoring, corrosion, and defect tracking, is a nice-to-have.

• Familiarity with regulatory and compliance reporting requirements for pipeline or asset integrity data is a nice-to-have.

• Experience migrating customers from legacy or spreadsheet-based systems to contemporary data platforms is a nice-to-have.


🏝️ Benefits

• Competitive salary and comprehensive benefits package.

• Opportunities for professional development and career advancement.

• Collaborative work environment with a focus on innovation.

• Flexible work hours and remote work options.

People also viewed

Data Elephant11 hours ago

Software Engineer – Data Products, AI Applications

CA flagCanada OnlyFreelanceData Engineer
ApplyView job
ICF19 hours ago

Senior Data Engineer, Scala

US flagVirginia OnlyFull-timeData Engineer$98.6k – $167.6k/year
ApplyView job
General Dynamics Information Technology1 day ago

Data Architect Principal, Analytics Pipeline

US flagUnited States OnlyFull-timeData Engineer$136k – $184k/year
ApplyView job
Logic20/20, Inc.1 day ago

Lead Geospatial Data Engineer

US flagWashington OnlyFull-timeData Engineer$173.1k – $179.9k/year
ApplyView job
Rox Partner1 day ago

Senior Data Engineer

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job
EVT1 day ago

Data Engineer – GCP

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers