
Lead Data Engineer
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Alabama, +1 more state.
• Establish the technical vision for a delivery pod focused on developing data products on Voyager, Protective’s Databricks lakehouse hosted on Azure.
• Oversee the comprehensive design of data products utilizing Raw/Prep/Prod (Bronze/Silver/Gold) medallion architecture.
• Manage dimensional design aspects, including grain, keys, Type 2 history, facts, bridges, and conformed dimensions.
• Define boundaries for data-cleaning, business logic, and consumer responsibilities across various layers.
• Collaborate with ML engineering to create contracted, versioned, and reproducible training and feature datasets.
• Oversee ODCS data contracts, quality standards, freshness expectations, consumer compatibility, and policies regarding breaking changes.
• Establish engineering standards for Python, SQL, dbt, testing, model structure, naming conventions, and repository practices.
• Conduct code reviews and ensure quality through testing, asset checks, observability, alerting, and CI/CD processes.
• Manage pipeline operations, including incident diagnosis, data triage, backfilling, tuning, on-call escalation, root-cause analysis, and runbook creation.
• Uphold regulated-carrier controls, ensure least-privilege access, enforce segregation of duties, manage changes, and maintain audit evidence.
• Create reusable frameworks, templates, and design patterns.
• Collaborate with the Product Owner and Scrum Master for story decomposition, refinement, and delivery preparedness.
• Mentor engineers through design reviews, pairing sessions, and code evaluations.
• Work closely with platform teams, DataOps/MLOps, data architecture, and governance teams.
• Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related discipline; equivalent practical experience is acceptable.
• Over 6 years of experience in building and operating production data pipelines and consumer-facing data models, from source ingestion to published data products.
• Strong practical skills in Python and SQL.
• Hands-on experience with Databricks or a similar Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance optimization.
• Extensive experience in dimensional modeling.
• Proven technical leadership capabilities.
• Experience managing data dependencies for other teams, including addressing breaking changes and production data incidents.
• Familiarity with orchestration tools such as Dagster, Databricks Workflows, Airflow, or comparable systems.
• Experience with Git-based collaborative development, code reviews, and CI/CD tools such as Azure DevOps or similar platforms.
• Familiarity with setting observability and SLA/SLO expectations for data relied upon by other teams.
• Ability to articulate trade-offs clearly to engineers, product owners, and business stakeholders.
• Databricks certification is preferred.
• Preferred experience with Unity Catalog, dbt, Dagster and Dagster Cloud, data contracts/ODCS, declarative Python ingestion frameworks, data quality and observability tools, MLOps, Azure and Azure DevOps, regulated industries, and AI-assisted development.
• Comprehensive health, dental, and vision insurance.
• Mental health benefits.
• Employee assistance program.
• Paid time off.
• Paid parental leave.
• Short-term disability coverage.
• Cultural observance days.
• Contributions to healthcare accounts.
• Pension plan.
• 401(k) plan with company matching.
• ProHealth Rewards platform offering cash rewards.
• Annual incentives based on individual and company performance.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.