Remotery

IT Engineer, Data Lakehouse

atContinentalRemoteIN flagIndiaFull-timeData EngineerMid-levelSenior

Posted Jul 18

This is a fully remote position, open to applicants in India.

📋 Description

• Design, develop, and manage scalable and maintainable data pipelines within the Azure Databricks environment.

• Create all technical artifacts as code, utilizing professional IDEs, with comprehensive version control and CI/CD automation.

• Facilitate data-driven decision-making in Supply Chain Management (SCM) by ensuring high levels of data availability, quality, and reliability.

• Develop data products and analytical assets by applying software engineering principles in close cooperation with business domains and functional IT.

• Employ stringent software engineering practices, such as modular design, test-driven development, and artifact reuse in all implementations.

• Global delivery presence, providing cross-functional data engineering support across SCM domains.

• Collaborate with business stakeholders, functional IT partners, product owners, architects, ML/AI engineers, and Power BI developers.

• Operate within an agile, product-team structure embedded in a large-scale Azure environment.

• Design scalable batch and streaming pipelines in Azure Databricks utilizing PySpark and/or Scala.

• Implement ingestion from structured and semi-structured sources (e.g., SAP, APIs, flat files).

• Construct bronze/silver/gold data layers in accordance with the defined lakehouse layering architecture and governance.

• Develop use-case driven dimensional models (star/snowflake schema) tailored to the needs of SCM.

• Ensure compatibility with reporting tools (e.g., Power BI) through curated data marts and semantic models.

• Implement enterprise-level data warehouse models (domain-driven 3NF models) for SCM data, closely collaborating with data engineers from other business domains.

• Develop and apply master data management strategies (e.g., Slowly Changing Dimensions).

• Create automated data validation tests utilizing frameworks.

• Monitor pipeline health, detect anomalies, and establish quality thresholds.

• Create data quality transparency by defining and implementing significant data quality rules with source system and business stakeholders, along with related reports.

• Develop and structure pipelines using modular, reusable code within a professional IDE.

• Implement test-driven development (TDD) principles with automated unit, integration, and validation tests.

• Integrate tests into CI/CD pipelines to facilitate fail-fast deployment strategies.

• Commit all artifacts to version control, ensuring peer reviews and CI/CD integration.

• Collaborate closely with Product Owners to refine user stories and establish acceptance criteria.

• Translate business requirements into data contracts and technical specifications.

• Engage in agile events such as sprint planning, reviews, and retrospectives.

• Document pipeline logic, data contracts, and technical decisions in markdown or auto-generated documents from code.

• Align designs with governance and metadata standards (e.g., Unity Catalog).

• Track lineage and audit trails through integrated tooling.

• Profile and optimize data transformation performance.

• Reduce job execution times and optimize cluster resource utilization.

• Refactor legacy pipelines or inefficient transformations to enhance scalability.


⛳️ Requirements

• Bachelor's degree in Computer Science, Data Engineering, Information Systems, or a related field.

• Certifications in software development and data engineering (e.g., Databricks DE Associate, Azure Data Engineer, or relevant DevOps certifications).

• 3–6 years of practical experience in data engineering roles within enterprise environments.

• Proven experience in building production-grade codebases in IDEs, complete with test coverage and version control.

• Demonstrated expertise in implementing complex data pipelines and contributing to full lifecycle data projects (from development to deployment).

• Experience in at least one business domain: SCM or a comparable area.

• While not mandatory, experience mentoring junior developers or leading implementation workstreams is advantageous.

• Experience collaborating with international teams across multiple time zones and cultures, preferably with teams in India, Germany, and the Philippines.


🏝️ Benefits

• Opportunities for training and development.

• Flexible and mobile working models.

• Sabbaticals and much more.

People also viewed

Railroad1912 hours ago

Senior Data Engineer – GCP, Python, Iceberg, Delta Lake, Kafka, Snowflake, Databricks

US flagUnited States OnlyFull-timeData Engineer$120k – $180k/year
ApplyView job
Livefront12 hours ago

Data Engineer

PE flagPeru OnlyFull-timeData Engineer
ApplyView job
GFT Technologies12 hours ago

Data Engineer, Mid-level

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
VIDA13 hours ago

Geospatial Data Engineer – Customer & AI Solutions

DE flagGermany OnlyFull-timeData Engineer
ApplyView job
albo14 hours ago

Data Engineer

MX flagMexico OnlyFull-timeData Engineer
ApplyView job
Leega14 hours ago

Engenheiro de Dados Pleno – AWS

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers