
Principal Data Engineer
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Design and develop dimensional models, including star schemas and snowflake schemas.
• Create and sustain semantic models to serve as the authoritative source for business reporting.
• Implement strategies for Slowly Changing Dimensions (SCD).
• Oversee master data engineering, which includes managing golden records, source-of-record authority, and cross-system identity resolution.
• Establish and uphold data modeling standards.
• Design and manage real-time and near-real-time data pipelines utilizing Kafka, Confluent Cloud, and Fabric Eventstreams.
• Minimize data staleness by enhancing scheduling, optimizing incremental loads, and orchestrating pipelines.
• Take charge of full-stack performance tuning, covering query optimization, partitioning, indexing, Delta table compaction, semantic model refreshes, and Direct Lake readiness.
• Implement operational engineering practices that include observability, alerting, SLAs, failure recovery, and capacity planning.
• Design and enforce controls for sensitive data, including financial information, PII, and HIPAA data.
• Develop robust, scalable, and observable ETL/ELT pipelines employing watermark-based incremental loads, CDC, batch, and streaming architectures.
• Ensure that pipelines are idempotent, recoverable, and production-ready.
• Act as the senior technical authority during code reviews.
• Convert business requirements into semantic models and report-layer artifacts.
• Serve as the main technical liaison for Power BI, AI/ML, and application development teams.
• Define and uphold data contracts related to schema stability, access patterns, and SLAs.
• Enhance the developer experience on the platform through improved discoverability, documentation, and onboarding processes.
• Establish data engineering standards, shape data strategy, mentor colleagues, and influence technical direction.
• Proficiency in Microsoft Fabric, encompassing Lakehouses, Notebooks, Dataflows Gen2, Event streams, Semantic Models, and Direct Lake mode.
• Expertise in Power BI report development, dataset/semantic model design, and DAX.
• Strong skills in SQL Server, Azure SQL, or Postgres query optimization, schema design, and stored procedures.
• Experience with Azure DevOps Git-based development workflows and CI/CD processes for data pipelines.
• Proven track record of utilizing AI coding assistants within a production engineering workflow.
• Ability to critically assess, edit, and enhance AI-generated code and artifacts.
• Understanding of the benefits AI brings to workflows and the associated risks.
• Familiarity with Azure Data Factory, Synapse Analytics, ADLS Gen2, and Event Hubs is preferred.
• Experience in a private equity-backed or multi-entity portfolio company setting is preferred.
• Exposure to MDM platforms such as Profisee, Semarchy, Ataccama, or similar is preferred.
• Experience with Confluent Cloud or Apache Kafka streaming ingestion is preferred.
• Familiarity with cross-tenant Azure/Fabric architecture is preferred.
• Background in business analysis, solutions architecture, or pre-sales engineering is preferred.
• Microsoft Fabric or Azure Data Engineer certifications are preferred.
• Health Care Plan (Medical, Dental & Vision)
• Retirement Plan (401k)
• Life Insurance (Basic, Voluntary & AD&D)
• Paid Time Off
• Short Term & Long-Term Disability
• Training & Development
• Wellness Resources
Progress Partners
FYUL
CarringtonCrisp
Get handpicked remote jobs straight to your inbox weekly.