
Data Engineer, Python, Spark
Posted Aug 25

Posted Aug 25
This is a fully remote position, open to applicants in Romania.
• Design, develop, and manage scalable batch, micro-batch, and streaming data pipelines utilizing Python, PySpark, and Azure Databricks.
• Integrate various data sources, both internal and external, including APIs, databases, files, event streams, and third-party feeds.
• Create ingestion solutions for structured, semi-structured, and unstructured data.
• Construct web scraping and data acquisition components when APIs are not available.
• Implement techniques such as pagination, throttling, retries, exponential backoff, checkpointing, schema evolution, and recovery patterns.
• Design idempotent processing, reprocessing, controlled backfills, and facilitate graceful recovery from partial failures.
• Develop and sustain batch and streaming patterns for market, weather, fundamental, and time-series data.
• Create modular, reusable, and testable components in Python and PySpark.
• Organize maintainable software projects and produce clean, well-documented code.
• Establish automated tests and integrate them into delivery pipelines.
• Conduct code reviews and advocate for engineering standards.
• Package reusable functionalities as Python modules or wheels.
• Diagnose complex issues across source systems, APIs, processing logic, infrastructure, and production environments.
• Design data models for analysts, traders, reporting solutions, and downstream data products.
• Implement Lakehouse and Medallion architecture across Bronze, Silver, and Gold layers.
• Design solutions addressing schema evolution, retention, lineage, and reproducible processing.
• Optimize data layouts, partitioning, joins, file sizes, caching, and Spark execution plans.
• Design, schedule, and manage workflows using Databricks Workflows and Astronomer.
• Implement dependency management, parameterization, environment configuration, and controlled promotion.
• Define runbooks and assist in diagnosis, recovery, and problem management.
• Monitor pipeline health, freshness, completeness, performance, and data quality.
• Utilize production-safe release patterns, including controlled rollouts, rollback, and validation.
• Manage production code through Git and adhere to pull-request/review practices.
• Build and maintain CI/CD pipelines using GitHub Actions and/or Azure DevOps.
• Deploy Databricks jobs, pipelines, and application artifacts using Databricks Asset Bundles or similar mechanisms.
• Provision and configure cloud and Databricks resources via Terraform.
• Apply automated validation, security scanning, and testing prior to production deployment.
• Contribute reusable pipeline templates, engineering standards, and platform automation.
• Engineer for availability, recoverability, scalability, and predictable operations.
• Optimize Spark workloads and select appropriate compute models and cluster configurations.
• Integrate cost-awareness into pipeline design, compute, scheduling, storage, and retention.
• Monitor resource consumption and work towards reducing processing times and cloud expenses.
• Implement automated data validation, schema checks, null checks, referential-integrity controls, and business quality rules.
• Identify and manage schema drift and changes in source data.
• Utilize Delta Lake and Unity Catalog for lineage, access control, metadata, and governance.
• Handle secrets and credentials securely through approved mechanisms and managed identities.
• Maintain technical documentation, metadata, and operational information.
• Collaborate with teams focused on data governance, architecture, security, and platform initiatives.
• Work alongside traders, analysts, data scientists, software engineers, product owners, and platform teams to translate requirements into technical solutions.
• Communicate design decisions, risks, dependencies, and trade-offs effectively.
• Contribute reusable components, templates, documentation, and engineering guidelines.
• Operate efficiently in a distributed, international, and cross-functional environment.
• Minimum of 5 years of professional experience.
• Preferably, experience in the energy sector and/or trading.
• Proficiency in operating fundamental power-price forecasting models for short- to mid-term trading.
• Capability to manage large volumes of data.
• Ability to work fully remotely within a collaborative environment alongside interdisciplinary teams.
• Strong communication skills and professionalism.
• Expertise in Python and PySpark.
• Experience with Azure Databricks.
• Familiarity with REST APIs, GraphQL, WebSocket, and gRPC interfaces.
• Knowledge of databases, files, event streams, and third-party data feeds.
• Proficiency in JSON, CSV, Parquet, and Delta formats.
• Experience in web scraping and data acquisition.
• Understanding of object-oriented and functional design principles.
• Familiarity with type hints, interfaces, and design patterns.
• Experience in unit, integration, contract, and data-quality testing.
• Proficiency in Git, GitHub Actions, and/or Azure DevOps.
• Experience with Databricks Asset Bundles or equivalent solutions.
• Knowledge of Terraform.
• Familiarity with Delta Lake and Unity Catalog.
• Understanding of secure secret-management mechanisms and managed identities.
• Medical insurance: Your health and that of your family are our top priority. You're fully covered.
• Holiday flat in Valencia.
• Gifts for special occasions, including Easter, Women's Day, Father's Day, and more.
• Anniversary gifts for 1st, 5th, and 10th work anniversaries.
• Team-building events.
• End-of-year celebrations.
• A day off on your birthday.
• Community and social initiatives.
• Unlimited access to thousands of courses through Udemy.
• German language courses.
• Masterclasses with experienced trainers.
• Investment in courses and training for personal and professional growth.
RR Donnelley
plotdesk
CmdScale GmbH
Colsubsidio
Get handpicked remote jobs straight to your inbox weekly.