
Data Architect – Mid-Level
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in United States.
• Develop the schema and unified data model for the platform, encompassing asset, inspection, incident, geospatial, and consequence data across various products.
• Implement Bronze, Silver, and Gold medallion architecture patterns and establish the semantic layer utilized by downstream models, reporting, and applications.
• Model data necessary for probabilistic risk models, specifying sources, transformation paths, and data quality expectations for essential input attributes.
• Design cross-product entity resolution for shared assets and entities.
• Define ingestion patterns for batch, streaming, and CDC sources, addressing schema evolution and slowly changing dimensions.
• Create architecture documentation, diagrams, data dictionaries, and decision records.
• Design the shared GIS layer, which includes pipeline centerlines, consequence-area polygons, right-of-way corridors, and parcel data.
• Model linear referencing and dynamic segmentation to align risk outcomes with centerline geometry.
• Define data structures for inspection data, repair history, weather history, soil characteristics, and satellite-derived information.
• Establish cataloging, classification, metadata, and lineage standards in Unity Catalog.
• Set data quality rules, validation gates, profiling, quarantine, and remediation processes.
• Define access control patterns, tenant isolation, and sensitive data handling protocols.
• Design comprehensive traceability from risk outputs to source records.
• Collaborate with data engineers to implement architectural designs and review production pipelines.
• Contribute to schema deployment, transformation logic, and performance optimization.
• Partner with application engineers and data scientists on serving layers, feature stores, and APIs.
• Support performance and cost optimization through partitioning, clustering, file layout, and compute sizing.
• Engage in design reviews and architecture alignment sessions.
• Translate domain and regulatory requirements into data structures and standards.
• Keep architecture documentation up to date.
• Deliver a well-documented and implemented data model that supports risk model inputs, ensures strong lineage, traceability, implementation clarity, and reliable entity resolution.
• 5–8 years of experience in data architecture, data modeling, or senior data engineering, with end-to-end ownership of a significant data model.
• Hands-on experience with Databricks, including Delta Lake, Unity Catalog, SQL Warehouses, and pipeline orchestration.
• Strong grasp of medallion and lakehouse architecture patterns, dimensional modeling, and semantic layer design.
• Advanced SQL skills and proficiency in Python or PySpark.
• Experience with at least one major cloud platform, preferably Azure.
• Practical experience with change data capture (CDC), slowly changing dimensions (SCD), schema evolution, and data quality frameworks.
• Working knowledge of data governance, including cataloging, lineage, classification, role-based access control (RBAC), and encryption.
• Excellent documentation and diagramming capabilities, with the ability to clearly communicate data models to both technical and non-technical stakeholders.
• Experience in designing geospatial data models and utilizing GIS tools, spatial joins, and linear referencing.
• Experience with multi-tenant platforms, data residency requirements, or regulated environments.
• Familiarity with metadata and lineage tools such as Unity Catalog, Microsoft Purview, or similar platforms.
• Experience modeling data specifically for machine learning or probabilistic risk models.
• Experience designing semantic layers for business intelligence (BI) applications.
• Databricks or cloud data certifications.
• Experience with AI-assisted coding tools such as Cursor or GitHub Copilot and/or agentic coding tools like Claude Code as part of a professional development workflow.
• Familiarity with pipeline or utility asset data, including inline inspection results, alignment sheets, facilities, and centerline geometry.
• Understanding of asset integrity concepts such as corrosion growth, defect tracking, and consequence-of-failure modeling.
• Awareness of regulatory reporting requirements related to pipeline integrity data.
• Experience migrating data from legacy desktop or spreadsheet-based systems.
• Comprehensive health insurance coverage.
• Flexible working hours and remote work options.
• Opportunities for professional development and certifications.
• Collaborative and inclusive company culture.
Data Elephant
ICF
General Dynamics Information Technology
Logic20/20, Inc.
Get handpicked remote jobs straight to your inbox weekly.