
Data Engineering Manager
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
• Provide technical leadership for the Data Engineering team through architectural guidance, code reviews, mentorship, and the implementation of engineering best practices.
• Mentor and nurture the growth of junior and mid-level Data Engineers.
• Establish engineering standards for software development, testing, CI/CD, documentation, observability, and operational excellence.
• Lead technical design discussions, assess architectural tradeoffs, and promote the adoption of modern engineering practices.
• Act as the technical lead for Penn Foster Group's Databricks Lakehouse platform.
• Design, construct, and maintain scalable enterprise data pipelines utilizing SQL, Python, Apache Spark, and Databricks.
• Define best practices for Databricks development encompassing Workflows, Repos, Jobs, notebooks, Python libraries, Git integration, cluster policies, SQL Warehouses, and deployment automation.
• Design and enhance Delta Lake architectures using Medallion patterns, Delta optimization, Liquid Clustering, and Photon.
• Create reusable ingestion frameworks for batch, streaming, CDC, and API-based integrations.
• Optimize Spark workloads for improved performance, scalability, reliability, and cost efficiency in the cloud.
• Develop trusted semantic data products, Genie Spaces, and Genie Ontologies for Databricks Genie.
• Establish best practices for semantic modeling, governed metrics, business metadata, and datasets that are ready for AI applications.
• Evaluate and implement new Databricks AI capabilities.
• Collaborate with Data Science and MLOps teams on machine learning, generative AI, and advanced analytics projects.
• Design scalable Lakehouse architectures that support analytics, reporting, AI, and operational data products.
• Implement enterprise governance with Unity Catalog, including data lineage, fine-grained security, metadata management, and access controls.
• Advocate for automated testing, monitoring, observability, data quality, and production reliability.
• Serve as the technical escalation point for complex production issues and lead root cause analysis efforts.
• Drive platform modernization initiatives while balancing delivery, scalability, maintainability, and operational excellence.
• Collaborate with Business Intelligence, Product, Platform Engineering, Security, and MLOps teams.
• Translate business requirements into technical solutions for enterprise reporting, analytics, AI, and strategic decision-making.
• Contribute to technical roadmaps, platform strategy, and the advancement of Data & Analytics capabilities.
• Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience.
• Over 8 years of experience in Data Engineering, Analytics Engineering, or Data Platform Engineering.
• A minimum of 3 years of experience leading technical initiatives and mentoring engineering teams.
• Proven success in developing engineers while promoting engineering excellence and technical standards.
• Strong communication and collaboration skills with both technical and business stakeholders.
• Expert knowledge of the Databricks Lakehouse Platform, including Apache Spark/PySpark, Delta Lake, Unity Catalog, Workflows & Jobs, Repos, SQL Warehouses, Cluster Policies, Serverless Compute, MLflow, Lakehouse Monitoring, Auto Loader, and Delta Live Tables/Lakeflow.
• Proficient in SQL and Python.
• Strong experience with BI tools such as Power BI, Tableau, or Business Objects.
• Extensive Microsoft Azure experience, encompassing ADLS Gen2, Microsoft Entra ID/Azure AD, RBAC, networking, and cloud security.
• Experience with implementing Medallion Architecture, dimensional modeling, and domain-oriented data products.
• In-depth understanding of Spark optimization techniques, including Adaptive Query Execution, partitioning, caching, Photon, Liquid Clustering, and Delta optimization.
• Familiarity with implementing CI/CD pipelines, Git-based development workflows, Infrastructure as Code, and automated testing.
• Strong understanding of data governance, metadata management, data quality, observability, security, and compliance.
• Candidates must complete a role-specific assessment as the initial step in the hiring process.
• Successful completion of applicable pre-employment screening requirements.
• Completion of federal employment eligibility verification through Form I-9.
• Preferred: experience with Databricks Genie, Genie Spaces, and Genie Ontologies.
• Preferred: experience in designing semantic models and AI-ready data products.
• Preferred: experience supporting enterprise BI platforms.
• Preferred: experience with machine learning platforms, MLOps, or Generative AI applications.
• Preferred: experience with dbt, Great Expectations, or similar tools.
• Preferred: Databricks Certified Data Engineer Professional and/or Microsoft Azure certifications.
• Medical insurance.
• Dental insurance.
• Vision insurance.
• Flexible spending.
• Generous paid time off.
• Sponsored volunteer opportunities.
• 401K with a company match.
• Free access to all of Penn Foster Group's online programs.
• Remote work arrangement.
• On-camera work environment with remote collaboration.
DTN
Veltris
Adaptive Biotechnologies Corp.
Unity
Get handpicked remote jobs straight to your inbox weekly.