
Databricks Migration Engineer
Posted Jul 14

Posted Jul 14
This is a fully remote position, open to applicants in Virginia.
• Lead the technical transition from legacy SQL Server stored procedures and ADF pipelines to Databricks Lakehouse (Delta Lake), ensuring adherence to best practice Lakehouse design.
• Convert traditional relational data warehousing concepts into scalable, distributed Lakehouse architectures (Bronze, Silver, Gold).
• Create robust, reusable ETL/ELT frameworks utilizing PySpark, Delta Live Tables (DLT), and Databricks Workflows.
• Structure and enhance the Gold Layer (dimensional models, star schemas) specifically aimed at optimizing Power BI performance.
• Enhance Databricks SQL Warehouses to accommodate high-concurrency, low-latency Power BI queries (DirectQuery and Import modes).
• Apply advanced optimization strategies, including Z-Ordering, data skipping, liquid clustering, and materialized views.
• Establish and enforce governance standards for cluster sizing, auto-scaling policies, and serverless SQL compute to balance performance and cost.
• Set up proactive monitoring dashboards to track Databricks Unit (DBU) consumption and pinpoint cost-saving opportunities.
• Develop best practices for partition strategies and file size management within Delta Lake.
• Create and implement a robust data security model utilizing Unity Catalog for centralized governance.
• Implement row-level and column-level security policies to ensure compliant data access for Power BI users and internal analysts.
• Align the Lakehouse security framework with existing enterprise Azure Active Directory (Microsoft Entra ID) and RBAC standards.
• Serve as the primary technical lead, conducting dedicated pair-programming sessions, workshops, and code reviews to guide the team in transitioning from SQL-centric to Spark-centric methodologies.
• Develop comprehensive technical documentation, including architecture diagrams, design patterns, and optimization playbooks.
• Establish a foundational knowledge transfer framework to ensure the internal team achieves full self-sufficiency post-migration.
• Communicate effectively in both verbal and written forms to diverse audiences, both technical and non-technical.
• Maintain an organized approach, completing tasks in a timely manner while paying close attention to details.
• Bachelor's degree or higher from an accredited institution in Computer Science, Engineering, or a related technical discipline.
• Over 5 years of experience in Data Engineering, Data System Development, or related fields.
• More than 5 years of experience with Cloud platforms (e.g., Azure, AWS, GCP).
• At least 1 year of experience leading complex, cross-functional data projects and technical teams.
• Proficient in Databricks Lakehouse, Apache Spark, Delta Lake, cloud-native databases, storage solutions, and distributed compute platforms.
• Familiarity with data warehousing, dimensional modeling, enterprise data lakes, incremental data loads, and metadata-driven ingestion and data quality frameworks using PySpark.
• Strong communication skills.
• Must be a US Citizen and able to obtain a Position of Public Trust Clearance.
• Must have lived in the US for the past 5 years and not have traveled outside the US for a total of 6 months or more during that time.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Retirement savings plan with company match.
• Opportunities for professional development and continuous learning.
• Flexible work arrangements to support work-life balance.
3M Consultancy
Get handpicked remote jobs straight to your inbox weekly.