
Data Engineer, Databricks, PySpark
Posted Aug 24

Posted Aug 24
This is a fully remote position, open to applicants in United States.
• Comprehend data-related dependencies, processes, system interactions, and concepts of source-of-truth.
• Identify and trace data flows across various source and target systems.
• Investigate issues across system interfaces, data pipelines, and downstream processes to pinpoint root causes.
• Utilize Databricks and SQL for data analysis, assumption validation, inconsistency detection, and resolution of incidents.
• Maintain and enhance reliable data pipelines through validation, reconciliation, monitoring, and thorough documentation.
• Operate within established version-control, deployment, change, and release protocols.
• Collaborate in an agile delivery environment, taking ownership of tasks from start to finish.
• Assist in crucial data engineering activities, ensuring continuity and effective knowledge transfer.
• Demonstrated experience in designing, developing, testing, and managing production-grade ETL/ELT pipelines.
• Practical experience with Databricks, Python/PySpark, and SQL.
• Experience in integrating and reconciling data across enterprise systems such as REST APIs, ITSM, ITFM, CMDB, IAM, ERP, monitoring, or similar.
• Familiarity with data validation, quality checks, reconciliation, monitoring, lineage, and schema evolution.
• Proficiency in Git-based version control, code reviews, automated testing, CI/CD, or similar practices.
• Experience with DEV/TEST/PROD environments and structured release/change management processes.
• Strong skills in root-cause analysis across interfaces, pipelines, and downstream systems.
• Proficient in English.
• Fully remote work arrangement.
Expleo Group
CodiLime
M3 USA
M3 USA
Get handpicked remote jobs straight to your inbox weekly.