
Senior Databricks Engineer – ML, Analytics
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in North Carolina, +3 more states.
• Collaborate directly with teams from the Department of Veterans Affairs to define use cases, strategize migrations, and implement effective solutions on Databricks.
• Evaluate on-premises workloads, which include SQL Server, SSIS packages, stored procedures, and scheduled jobs, and modify them for Databricks patterns in Azure Commercial.
• Develop and optimize pipelines utilizing PySpark, Spark SQL, Delta Lake, Lakeflow, and Azure Data Factory.
• Employ Unity Catalog for access management, data lineage, and sharing to oversee customer data governance.
• Educate customers through workshops, collaborative sessions, and comprehensive documentation.
• Create reusable templates and code packages using Git, CI/CD practices, and Databricks Asset Bundles.
• AI/ML emphasis: Construct end-to-end machine learning workflows encompassing feature engineering, model training, evaluation, MLflow tracking, Unity Catalog model registry, and model serving.
• AI/ML emphasis: Assist data scientists in transitioning models into governed and monitored production environments.
• AI/ML emphasis: Support generative AI initiatives, including retrieval-augmented generation and agent frameworks.
• Analytics emphasis: Transform SSIS and stored procedure ETL processes into Databricks pipelines with reconciliation.
• Analytics emphasis: Substitute on-premises reporting with Databricks SQL, AI/BI dashboards, and Power BI.
• Analytics emphasis: Disseminate reporting data from Databricks SQL warehouses, including semantic models and optimizing query performance.
• Over 7 years of experience in data engineering, with at least 2 years dedicated to building production workloads on Databricks.
• Strong expertise in Python, SQL, PySpark, and Spark SQL.
• Practical experience with Delta Lake, medallion architecture design, and Spark performance optimization.
• Proficiency in Unity Catalog: handling catalogs, grants, lineage, and ensuring row- and column-level security.
• Experience in migrating on-premises workloads (e.g., SQL Server, SSIS, stored procedures) to cloud environments.
• Familiarity with Azure services, such as ADLS Gen2, Azure Data Factory, Entra ID, and Key Vault.
• Knowledge of Git, CI/CD processes, and infrastructure as code (e.g., Databricks Asset Bundles, Terraform).
• Experience in customer-facing consulting or professional services, including conducting requirement sessions, workshops, and training.
• Ability to communicate clearly, both written and verbally, with audiences that are technical and non-technical.
• Bachelor’s degree in computer science, information systems, or a related discipline.
• Must be a U.S. citizen.
• Must be able to obtain a public trust clearance.
• Must be eligible to work in the U.S.
• AI/ML emphasis: Experience in training, evaluating, and deploying models using MLflow and popular frameworks (e.g., scikit-learn, XGBoost, PyTorch).
• Analytics emphasis: Experience with Databricks SQL and Power BI data modeling (including star schemas, semantic models, DAX).
• Preferred: Databricks certifications; experience with healthcare data and HIPAA compliance; familiarity with Azure Government or FedRAMP High; experience in upgrading from Hive metastore to Unity Catalog; knowledge of Mosaic AI Vector Search, Agent Framework, or model monitoring; experience with Microsoft Fabric, Pyramid Analytics, or Genie.
• Competitive salary.
• Generous annual leave and paid holidays.
• Comprehensive group health and dental insurance plans.
• 401(k) plan with company matching.
• Life insurance and AD&D coverage.
• Continuous training and professional development opportunities.
Thrive Market
Shield AI
Weekday (YC W21)
Roadpass Digital
Get handpicked remote jobs straight to your inbox weekly.