Remotery

Senior Data Engineer – Databricks

Posted Jul 31

This is a fully remote position, open to applicants in Argentina, +5 more countries.

📋 Description

• Take charge of Databricks production support for the predictive data platform, focusing on monitoring, alerting, and incident response for all production data flows.

• Oversee and report on SLA performance metrics related to data pipeline delivery, ensuring transparency regarding platform health and accountability among both internal and external stakeholders.

• Discover and implement optimizations for data pipelines that minimize Databricks compute costs, enhance throughput, and shorten processing windows, while tracking the effects through measurable KPIs.

• Transition legacy ETL/ELT pipelines to Databricks, developing automation tools to lessen manual intervention and guarantee uninterrupted data delivery during the migration.

• Facilitate onboarding for new customers by provisioning, validating, and securing tenant data pipelines that ensure reliable and isolated data from the outset.

• Design and construct high-performance Databricks pipelines that efficiently ingest, transform, and serve ERP and CRM data on a large scale across both Azure and AWS environments.

• Manage the Delta Lake architecture, which includes schema design, partitioning strategies, data quality enforcement, and incremental processing patterns.

• Uphold data security best practices across Databricks environments, encompassing role-based access control, secrets management, and compliance obligations for enterprise CRM and ERP data.

• Implement data quality monitoring and observability throughout pipeline health and ML model inputs, ensuring data integrity that directly supports the accuracy of model predictions.

• Enforce multi-tenant data isolation patterns to guarantee secure and dependable data delivery across enterprise clients.

• Collaborate with the Enterprise Architecture team to ensure seamless integration of data pipelines within the broader product ecosystem.

• Provide support for a globally distributed operation through on-call rotations and after-hours incident response, meeting SLAs across various time zones.

• Maintain technical documentation, runbooks, and architectural decision records, contributing to team knowledge sharing and operational preparedness in on-call and incident response situations.

• Implement CI/CD best practices in data pipeline development, which includes version control, automated testing, and deployment tools to guarantee reliable and repeatable pipeline delivery.


⛳️ Requirements

• Minimum of 4 years of experience in data engineering.

• At least 2 years of experience with Databricks or the Apache Spark ecosystem in Azure and/or AWS environments.

• Expertise in PySpark, SQL, and Python with a proven history of building and managing production-grade pipelines under SLA constraints.

• Practical experience with Delta Lake, including schema evolution, ACID transactions, optimize/vacuum lifecycle, and both incremental and streaming processing methods.

• Proven hands-on experience with optimizing pipeline performance and compute efficiency in production Databricks settings.

• Strong working knowledge of PostgreSQL, including query optimization, schema design, and its use as a source or sink in production data pipelines.

• Experience in supporting and maintaining legacy ETL tools (SSIS, Informatica, custom Python/SQL pipelines, or similar) in a production environment.

• Familiarity with large-scale multi-tenant architectures, emphasizing tenant isolation, individual tenant performance, and data privacy, including navigating tools and platforms that typically assume single-tenant configurations.

• Demonstrated ability to work collaboratively across Data Science, Product, and Infrastructure teams, managing end-to-end delivery in a cross-functional setting.

• Strong understanding of data governance, security, and compliance principles, including access control, data privacy, and the protection of sensitive enterprise data across multi-tenant environments.


🏝️ Benefits

• Competitive salary and performance-based bonuses.

• Comprehensive health, dental, and vision insurance.

• Flexible working hours and remote work options.

• Opportunities for professional development and continuous learning.

• Collaborative work environment with a focus on innovation.

People also viewed

Railroad195 hours ago

Senior Data Engineer – GCP, Python, Iceberg, Delta Lake, Kafka, Snowflake, Databricks

US flagUnited States OnlyFull-timeData Engineer$120k – $180k/year
ApplyView job
Livefront6 hours ago

Data Engineer

PE flagPeru OnlyFull-timeData Engineer
ApplyView job
GFT Technologies6 hours ago

Data Engineer, Mid-level

BR flagBrazil OnlyFull-timeData Engineer
ApplyView job
VIDA7 hours ago

Geospatial Data Engineer – Customer & AI Solutions

DE flagGermany OnlyFull-timeData Engineer
ApplyView job
albo7 hours ago

Data Engineer

MX flagMexico OnlyFull-timeData Engineer
ApplyView job
Leega7 hours ago

Engenheiro de Dados Pleno – AWS

BR flagBrazil OnlyFreelanceData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers