
Senior Data Engineer
Posted 10 hours ago

Posted 10 hours ago
This is a fully remote position, open to applicants in United States.
• Design, develop, and enhance production-grade ETL/ELT pipelines spanning across Bronze, Silver, and Gold medallion layers.
• Take responsibility for optimizing Databricks platform performance and ensuring cost efficiency, which includes cluster/job sizing, Photon, partitioning, Z-ordering, Liquid Clustering, Auto Loader, and DBU governance.
• Architect and implement data models that support Microsoft Fabric/Power BI and downstream analytics applications.
• Build and sustain CI/CD pipelines for Databricks assets utilizing Asset Bundles, Repos, and Git-based deployment strategies.
• Integrate platform pipelines with middleware, source systems, and cloud-native services.
• Manage and enhance Unity Catalog, focusing on access control, lineage, row/column-level security, and workspace-catalog bindings.
• Execute data quality, observability, reliability, monitoring, alerting, and SLA management practices.
• Collaborate with security, compliance, and audit teams to uphold SOX ITGC alignment.
• Diagnose production pipeline issues, conduct root-cause analysis, and implement preventive measures.
• Create architectural documentation, procedures, and operational runbooks.
• Assess, pilot, and productionize Databricks-native AI/ML capabilities including Genie, MLflow, Feature Store, and Model Serving.
• Support vector search and RAG patterns on governed Unity Catalog data.
• Prepare curated ML-ready Gold-layer datasets and feature pipelines in collaboration with data science and analytics stakeholders.
• Monitor the Databricks roadmap and advocate for the adoption of Lakehouse AI, Mosaic AI, and agent framework.
• Establish guardrails and human-in-the-loop controls for AI-assisted and agentic engineering workflows.
• Mentor mid-level data engineers on Databricks best practices, code quality, and architecture.
• Convert business requirements into technical specifications for BI and AI solutions.
• Work closely with product, security, business, data science, and analytics stakeholders.
• Contribute to maintaining data governance, security, and privacy standards.
• Bachelor’s degree in computer science, information systems, engineering, statistics, or a related field, or equivalent professional experience.
• Over 7 years of experience as a data engineer.
• More than 3 years specifically focused on architecting and managing production workloads on Databricks.
• Strong grasp of data governance, data security, and access control best practices.
• Familiarity with Agile development methodologies, CI/CD automation, and Test-Driven Development.
• Exceptional problem-solving abilities and capacity to lead technical troubleshooting autonomously.
• Strong written and verbal communication skills, with the capability to explain technical concepts to non-technical audiences.
• Proven experience managing a lakehouse/medallion architecture at scale, including data modeling for BI usage.
• Experience working in a governed or regulated environment (SOX, HIPAA, or similar) with formal change control and access governance.
• In-depth, hands-on expertise with the Databricks Lakehouse Platform: Delta Lake, Unity Catalog, Delta Live Tables/Lakeflow, Workflows, and optimization of clusters/jobs (Photon, Auto Loader, Liquid Clustering).
• Advanced SQL and proficient Python (PySpark) development skills; familiarity with Scala is a plus.
• Experience with Databricks Asset Bundles, Repos, and CI/CD for lakehouse deployments.
• Working knowledge of cloud data services (Azure preferred — ADLS, Azure SQL, Synapse/Fabric; AWS/GCP equivalents are also acceptable) and cloud migration strategies.
• Hands-on experience with MLflow for experiment tracking, model registry, and lifecycle management.
• Working knowledge of Databricks AI/ML capabilities — Feature Store, Model Serving, Genie, Mosaic AI, or similar lakehouse ML tools.
• Familiarity with vector search, embeddings, or RAG architectures.
• Comfort in collaborating with data science teams on ML-ready data pipelines.
• Databricks certifications are highly preferred.
• Knowledge of Microsoft Fabric/Power BI and enterprise middleware such as Boomi is a plus.
• Discounted child care benefits.
• Medical, dental, and vision coverage for employees’ families and pets.
• Employee assistance programs that support mental health and personal development.
• Health and wellness initiatives.
• Paid time off.
• Discounts on work-related necessities, such as cell phones.
Deckers Brands
praxipal
Revecore
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.