
Data Engineer – Mid Level
Posted Sep 10

Posted Sep 10
This is a fully remote position, open to applicants in India.
• Design, develop, and enhance batch and streaming data ingestion pipelines across AWS, Azure, and GCP.
• Create pipelines utilizing Databricks Workflows, Apache Spark/PySpark, SQL, Delta Live Tables, and Databricks Lakeflow components.
• Execute Bronze–Silver–Gold medallion architecture, CDC, SCD Type 1 and Type 2, schema evolution, validation, reconciliation, and data-quality rules.
• Set up and manage Delta Lake storage frameworks, tables, schemas, partitions, and optimization processes.
• Implement OPTIMIZE, Z-ORDER, VACUUM, partitioning, and file management strategies.
• Support Unity Catalog metadata, including cataloging, lineage, governance, and integration with Microsoft Purview.
• Integrate Amazon S3, Azure Storage, and Google Cloud Storage with Databricks.
• Establish data-quality checks, profiling, validation, monitoring, RBAC policies, security controls, and data classification tags.
• Construct, schedule, oversee, and maintain production workflows using Databricks Workflows, Delta Live Tables, Azure Data Factory, and other approved tools.
• Contribute to CI/CD pipelines, automated testing, deployment, environment management, and DEV–QA–PROD promotion.
• Monitor production pipelines, troubleshoot failures, investigate root causes, support recovery, and optimize performance.
• Collaborate with the Senior Data Architect, Data Scientists, ML Engineers, Analysts, Product teams, and engineering stakeholders.
• Engage in architecture reviews, technical design discussions, coding reviews, and engineering standards meetings.
• Document pipelines, data flows, data dictionaries, transformation logic, data-quality rules, test cases, job schedules, and operational procedures.
• Convert architecture designs and technical standards into reliable, production-ready pipelines and platform capabilities.
• 3–5 years of experience in Data Engineering, ETL development, or cloud data platform engineering.
• Practical experience with Databricks, Apache Spark, PySpark, or other distributed data processing technologies.
• Strong expertise in SQL, encompassing structured data transformation, joins, aggregations, and performance-aware query development.
• Familiarity with at least one major cloud platform; Microsoft Azure preferred, with AWS and/or GCP also advantageous.
• Knowledge of data modeling, data quality, schema evolution, data validation, pipeline monitoring, and troubleshooting.
• Basic understanding of RBAC, encryption, credential and secret management, and secure access to cloud and data platform resources.
• Working knowledge of Delta Lake, medallion architecture, and modern lakehouse best practices is preferred.
• Experience with metadata, cataloging, and governance platforms such as Unity Catalog, Microsoft Purview, or AWS Glue Data Catalog is preferred.
• Familiarity with workflow orchestration and scheduling technologies such as Azure Data Factory, Databricks Workflows, Apache Airflow, or similar frameworks is preferred.
• Experience with Git-based development, CI/CD, and DevOps practices is preferred.
• Knowledge or experience in geospatial/GIS data, BI semantic layers, particularly Power BI, or data preparation for AI/ML workloads is preferred.
• Relevant cloud or Databricks certifications are preferred.
• Understanding of asset integrity management concepts is a plus.
• Experience with oil & gas, utility, infrastructure, or pipeline asset data is a plus.
• Familiarity with regulatory, compliance, and audit reporting requirements is a plus.
• Competitive compensation package based on experience and qualifications.
• Medical, Dental, and Vision Insurance.
• 401(k) Plan with Company Match.
• Generous Paid Time Off (PTO).
• Company-Paid Holidays.
• Flexible Work Options / work-from-home opportunities, depending on role and business needs.
• On-Call Compensation for eligible on-call shifts.
Data Elephant
ICF
General Dynamics Information Technology
Logic20/20, Inc.
Get handpicked remote jobs straight to your inbox weekly.