
Data Engineer – Mid Level
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in United States.
• Develop and sustain data ingestion pipelines for both structured and semi-structured sources, such as GIS, inline inspection data, SCADA, maintenance systems, and enterprise record systems.
• Execute batch and streaming data ingestion utilizing Databricks Workflows, Spark, PySpark, SQL, and declarative pipeline tools.
• Employ Bronze, Silver, and Gold medallion architecture patterns for data transformation, standardization, and enrichment.
• Implement change data capture, manage slowly changing dimensions, support schema evolution, and set data-validation rules.
• Normalize third-party and public data feeds, including historical weather data, soil characteristics, satellite-derived information, and one-call ticket data.
• Construct AI-assisted pipelines for the normalization of units, schemas, and semantics across disparate customer data.
• Develop automated data-quality repair workflows with provenance for synthesized values.
• Productionize document-extraction pipelines in collaboration with data scientists.
• Establish human-in-the-loop review and exception workflows for low-confidence extractions.
• Configure and manage Delta Lake tables, including partitioning strategies and optimization routines.
• Implement metadata, lineage, and cataloging standards through Unity Catalog.
• Build and maintain connectors to customer systems of record with customizable refresh schedules.
• Support geospatial data processing, including spatial joins and alignment with pipeline centerline geometry.
• Implement data-quality tests, profiling, drift monitoring, access controls, security protocols, classification tags, and regulatory traceability.
• Create, schedule, and oversee workflows; manage alerting and handling of pipeline failures.
• Contribute to CI/CD processes for pipeline code, encompassing version control, automated testing, and environment promotion.
• Resolve production incidents, recover failed pipeline executions, and optimize performance and infrastructure expenses.
• Translate Data Architect designs into practical implementations and identify design gaps or ambiguities.
• Engage in architecture, design, and code reviews.
• Document pipelines, transformation logic, data dictionaries, job schedules, operational procedures, and runbooks.
• Deliver dependable Bronze, Silver, and Gold data flows, reduce the manual onboarding effort for customer data, maintain high data-quality pass rates, comply with governance standards, minimize incidents, and collaborate effectively with architects, data scientists, application engineers, and cross-functional partners.
• 3–5 years of experience in data engineering, ETL development, or cloud data platform engineering.
• Practical experience with Databricks, Spark, PySpark, or similar distributed data-processing technologies.
• Proficient SQL skills and experience in structured data transformation.
• Familiarity with at least one major cloud platform; Azure experience is preferred.
• Knowledge of data modeling, data-quality practices, schema evolution, and pipeline troubleshooting.
• Experience with workflow orchestration and scheduling frameworks.
• Understanding of essential data-security practices, including access control, encryption, and credential management.
• Experience with Git-based development and comfort collaborating within a code-reviewed engineering team.
• Preferred experience with Delta Lake, medallion architecture, and lakehouse engineering best practices.
• Experience with Unity Catalog, Microsoft Purview, or similar metadata and data-lineage tools is preferred.
• Experience building pipelines for the ingestion of unstructured or semi-structured documents is preferred.
• Familiarity with geospatial data processing and common GIS data formats is preferred.
• CI/CD and DevOps experience for data workloads, including infrastructure as code (IaC), is preferred.
• Experience in preparing and transforming data specifically for machine learning or probabilistic model consumption is preferred.
• Cloud or Databricks certifications are preferred.
• Experience using AI-assisted coding tools such as Cursor or GitHub Copilot and/or agentic coding tools like Claude Code within a professional development workflow is preferred.
• Experience integrating oil and gas or utility asset data, including pipelines, facilities, and GIS assets, is a nice-to-have.
• Understanding of asset integrity concepts, including inspection data, risk scoring, corrosion, and defect tracking, is a nice-to-have.
• Familiarity with regulatory and compliance reporting requirements for pipeline or asset integrity data is a nice-to-have.
• Experience migrating customers from legacy or spreadsheet-based systems to contemporary data platforms is a nice-to-have.
• Competitive salary and comprehensive benefits package.
• Opportunities for professional development and career advancement.
• Collaborative work environment with a focus on innovation.
• Flexible work hours and remote work options.
Data Elephant
ICF
General Dynamics Information Technology
Logic20/20, Inc.
Get handpicked remote jobs straight to your inbox weekly.