
Senior Data Engineer – Databricks
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in Argentina, +5 more countries.
• Take charge of Databricks production support for the predictive data platform, focusing on monitoring, alerting, and incident response for all production data flows.
• Oversee and report on SLA performance metrics related to data pipeline delivery, ensuring transparency regarding platform health and accountability among both internal and external stakeholders.
• Discover and implement optimizations for data pipelines that minimize Databricks compute costs, enhance throughput, and shorten processing windows, while tracking the effects through measurable KPIs.
• Transition legacy ETL/ELT pipelines to Databricks, developing automation tools to lessen manual intervention and guarantee uninterrupted data delivery during the migration.
• Facilitate onboarding for new customers by provisioning, validating, and securing tenant data pipelines that ensure reliable and isolated data from the outset.
• Design and construct high-performance Databricks pipelines that efficiently ingest, transform, and serve ERP and CRM data on a large scale across both Azure and AWS environments.
• Manage the Delta Lake architecture, which includes schema design, partitioning strategies, data quality enforcement, and incremental processing patterns.
• Uphold data security best practices across Databricks environments, encompassing role-based access control, secrets management, and compliance obligations for enterprise CRM and ERP data.
• Implement data quality monitoring and observability throughout pipeline health and ML model inputs, ensuring data integrity that directly supports the accuracy of model predictions.
• Enforce multi-tenant data isolation patterns to guarantee secure and dependable data delivery across enterprise clients.
• Collaborate with the Enterprise Architecture team to ensure seamless integration of data pipelines within the broader product ecosystem.
• Provide support for a globally distributed operation through on-call rotations and after-hours incident response, meeting SLAs across various time zones.
• Maintain technical documentation, runbooks, and architectural decision records, contributing to team knowledge sharing and operational preparedness in on-call and incident response situations.
• Implement CI/CD best practices in data pipeline development, which includes version control, automated testing, and deployment tools to guarantee reliable and repeatable pipeline delivery.
• Minimum of 4 years of experience in data engineering.
• At least 2 years of experience with Databricks or the Apache Spark ecosystem in Azure and/or AWS environments.
• Expertise in PySpark, SQL, and Python with a proven history of building and managing production-grade pipelines under SLA constraints.
• Practical experience with Delta Lake, including schema evolution, ACID transactions, optimize/vacuum lifecycle, and both incremental and streaming processing methods.
• Proven hands-on experience with optimizing pipeline performance and compute efficiency in production Databricks settings.
• Strong working knowledge of PostgreSQL, including query optimization, schema design, and its use as a source or sink in production data pipelines.
• Experience in supporting and maintaining legacy ETL tools (SSIS, Informatica, custom Python/SQL pipelines, or similar) in a production environment.
• Familiarity with large-scale multi-tenant architectures, emphasizing tenant isolation, individual tenant performance, and data privacy, including navigating tools and platforms that typically assume single-tenant configurations.
• Demonstrated ability to work collaboratively across Data Science, Product, and Infrastructure teams, managing end-to-end delivery in a cross-functional setting.
• Strong understanding of data governance, security, and compliance principles, including access control, data privacy, and the protection of sensitive enterprise data across multi-tenant environments.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and remote work options.
• Opportunities for professional development and continuous learning.
• Collaborative work environment with a focus on innovation.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.