
Data Platform Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Romania, +1 more country.
• Design and manage change-data-capture pipelines from PostgreSQL to Azure utilizing Kafka Connect and Debezium.
• Set up, deploy, and scale connectors end-to-end, encompassing connector configuration, task management, offsets, schema history, and snapshot strategies.
• Operate pipelines as stateful workloads on Kubernetes AKS, focusing on configurations, secrets, networking, and resource optimization.
• Oversee and resolve issues on the platform in production, addressing connector failures, task rebalances, restarts, throughput, backpressure, message-size limits, retries, and recovery.
• Automate the platform using Python through configuration-driven onboarding of data sources, orchestration of pipelines, monitoring, alerting, recovery workflows, and automated testing.
• Integrate CDC streams with Azure Event Hubs, ADLS, Azure PostgreSQL, ADF, and Databricks.
• Manage platform infrastructure as code to ensure reproducibility of environments and reviewability of changes.
• Implement data protection measures for sensitive data traversing the pipelines, including masking, hashing, access control, and retention policies.
• Robust commercial experience as a data or platform engineer, with practical exposure to streaming or CDC pipelines beyond just batch reporting.
• In-depth knowledge of Kafka, including topics, partitions, offsets, consumer groups, and delivery semantics.
• Experience with at least one independently run Kafka Connect deployment.
• Proficient SQL and PostgreSQL skills.
• Familiarity with WAL, logical replication, replication slots, and replication lag.
• Comprehensive understanding of CDC concepts, including initial snapshots, inserts, updates, deletes, event ordering, at-least-once delivery, and schema evolution.
• Confident in using Python for automation and tooling purposes.
• Practical experience with Azure data services like Event Hubs, ADLS, or Azure PostgreSQL.
• Comfortable working with Kubernetes, including deploying workloads, managing configurations and secrets, reading logs, and debugging malfunctioning pods.
• Ability to debug active pipelines through metrics and logs.
• Production experience specifically with Debezium.
• Experience managing stateful workloads on AKS, including StatefulSets, consistent worker identity, and resource tuning under load.
• Familiarity with infrastructure-as-code and CI/CD practices for data platform components, utilizing Terraform, Bicep, or similar tools.
• Hands-on experience with Databricks and ADF at a production scale.
• Experience in implementing data protection controls for sensitive data, including masking, hashing, access control, and retention policies.
• Flexible and remote working arrangements.
• Opportunities to work on international projects with prominent global clients.
• Travel opportunities related to international projects.
• A non-corporate work environment.
• Access to language classes.
• Participation in knowledge-sharing initiatives.
• Access to private healthcare and life insurance.
• Multisport card access.
Mirantis
Get handpicked remote jobs straight to your inbox weekly.