
Data Platform Engineer
Posted Aug 21

Posted Aug 21
This is a fully remote position, open to applicants in Poland.
• Design and maintain change-data-capture pipelines transferring data from PostgreSQL to Azure utilizing Kafka Connect and Debezium.
• Set up, deploy, and scale connectors comprehensively, involving connector configuration, task management, offsets, schema history, and snapshot strategies.
• Operate pipelines as stateful workloads on Kubernetes (AKS), addressing configuration, secrets management, networking, and resource optimization.
• Oversee and resolve issues within the production platform, including connector errors, task rebalances, restarts, throughput challenges, backpressure, message-size limitations, retries, and recovery processes.
• Automate platform functions using Python, focusing on configuration-driven onboarding, pipeline orchestration, monitoring and alerting, recovery workflows, and automated testing.
• Integrate CDC streams with Azure Event Hubs, ADLS, Azure PostgreSQL, ADF, and Databricks.
• Manage platform infrastructure as code to ensure environments are replicable and alterations are subject to review.
• Implement data protection measures for sensitive information traversing pipelines, encompassing masking, hashing, access controls, and retention policies.
• Proven commercial experience as a data or platform engineer, with direct involvement in streaming or CDC pipelines, as opposed to solely batch reporting.
• Practical knowledge of Kafka — including topics, partitions, offsets, consumer groups, and delivery semantics — demonstrated by at least one Kafka Connect deployment you managed independently.
• Proficient SQL and PostgreSQL capabilities, with a solid understanding of WAL, logical replication, replication slots, and replication lag.
• Familiarity with CDC principles: initial snapshots, inserts, updates, deletes, event ordering, at-least-once delivery, and schema evolution.
• Competent in Python for automation and tooling purposes — including orchestration, monitoring, recovery scripts, and automated testing.
• Practical experience with Azure data services, such as Event Hubs, ADLS, or Azure PostgreSQL.
• Comfortable using Kubernetes: deploying workloads, managing configurations and secrets, reading logs, and troubleshooting failing pods.
• Ability to analyze a running pipeline through metrics and logs — distinguishing throughput issues caused by backpressure, retries, or actual connector failures.
• Production experience with Debezium, specifically concerning snapshot strategies for large tables, schema history recovery, offset loss, and reconnecting connectors post-failure.
• Experience managing stateful workloads on AKS: working with StatefulSets, ensuring stable worker identities, and resource optimization under load.
• Familiarity with infrastructure-as-code and CI/CD practices for data platform components (Terraform, Bicep, or similar).
• Practical experience with Databricks and ADF at a production scale.
• Hands-on experience implementing data protection controls for sensitive data, including masking, hashing, access controls, and retention policies.
• Flexible employment options and the possibility of remote work.
• International projects with top-tier global clients.
• Opportunities for international travel.
• A non-corporate work environment.
• Language courses available.
• Internal and external training programs.
• Private healthcare and insurance coverage.
• Multisport card access.
• Initiatives focused on well-being.
Cypher Consulting Europe S.L.
Sigma Software Group
Coinbase
Get handpicked remote jobs straight to your inbox weekly.