
Data Platform Engineer
Posted Sep 2

Posted Sep 2
This is a fully remote position, open to applicants in Romania, +1 more country.
• Develop and manage change-data-capture pipelines from PostgreSQL to Azure utilizing Kafka Connect and Debezium.
• Set up, deploy, and scale connectors from start to finish, encompassing connector configuration, task management, offsets, schema history, and snapshot strategies.
• Operate pipelines as stateful workloads on Kubernetes (AKS), addressing configuration, secrets, networking, and resource optimization.
• Oversee and troubleshoot the production platform, including handling connector failures, task rebalances, restarts, throughput, backpressure, message-size constraints, retries, and recovery procedures.
• Automate platform operations in Python through configuration-led onboarding, pipeline orchestration, monitoring and alerting, recovery workflows, and automated testing.
• Integrate CDC streams with Azure Event Hubs, ADLS, Azure PostgreSQL, ADF, and Databricks.
• Manage platform infrastructure as code to ensure reproducibility of environments and reviewability of changes.
• Implement data protection measures for sensitive information traversing pipelines, including masking, hashing, access control, and retention policies.
• Strong commercial experience as a data or platform engineer, with practical involvement in streaming or CDC pipelines rather than solely in batch reporting.
• Proficient knowledge of Kafka, including topics, partitions, offsets, consumer groups, and delivery semantics, with at least one Kafka Connect deployment personally managed.
• Advanced SQL and PostgreSQL expertise, with a solid understanding of WAL, logical replication, replication slots, and replication lag.
• Familiarity with CDC concepts such as initial snapshots, inserts, updates, deletes, event ordering, at-least-once delivery, and schema evolution.
• Proficient in Python for automation and tooling, including orchestration, monitoring, recovery scripts, and automated tests.
• Practical experience with Azure data services, such as Event Hubs, ADLS, or Azure PostgreSQL.
• Comfortable utilizing Kubernetes as a user, including deploying workloads, managing configuration and secrets, reading logs, and debugging faulty pods.
• Ability to analyze a running pipeline using metrics and logs to differentiate between throughput issues and problems related to backpressure, retries, or actual connector failures.
• Production experience with Debezium, particularly in managing snapshot strategies for large tables, schema history recovery, offset loss, and re-establishing connectors after failures.
• Experience in operating stateful workloads on AKS, specifically with StatefulSets, stable worker identity, and resource tuning during high load periods.
• Familiarity with infrastructure-as-code and CI/CD practices for data platform components, such as Terraform or Bicep.
• Hands-on experience with Databricks and ADF at a production level.
• Experience in applying data protection controls for sensitive information, including masking, hashing, access control, and retention policies.
• Flexible work arrangements and remote employment options.
• Opportunities to work on international projects with leading global clients.
• Possibility of international travel for business purposes.
• A non-corporate work environment.
• Language learning opportunities.
• Access to internal and external training programs.
• Private healthcare and insurance coverage.
• Multisport card benefits.
• Initiatives focused on employee well-being.
Mirantis
Get handpicked remote jobs straight to your inbox weekly.