
Senior Data Developer β Databricks
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in Brazil.
β’ We are in search of a Senior Data Developer (Databricks) to become a vital member of our team, taking charge of a crucial data platform that underpins business reporting and analytics for our client.
β’ This role offers the chance to delve deeply into the Databricks ecosystem β from initial data ingestion to a fully modeled dimensional layer β while serving as the technical foundation on which the rest of the team depends for platform expertise.
β’ The position combines hands-on execution with technical leadership: you will manage and enhance notebooks throughout the complete medallion architecture, design and expand dimensional modeling artifacts that drive downstream business intelligence, and act as the primary reviewer and technical go-to for the team.
β’ You will function both strategically β by suggesting architectural and process enhancements β and operationally β by troubleshooting production jobs, deploying across environments, and ensuring that documentation remains up to date.
β’ Delivery & Continuous Improvement: Manage assigned tickets and bugs while continuously seeking improvement opportunities beyond the immediate task β proactively identifying technical debt, refactoring possibilities, and automation gaps instead of limiting contributions to what is assigned.
β’ Data Pipeline Ownership: Manage and enhance notebooks across the Stage, Bronze, Silver, and Gold layers, from data ingestion through the dimensional model utilized by business intelligence reporting tools.
β’ Dimensional Modeling: Design and implement dimensional modeling artifacts β including facts, dimensions, and slowly changing dimensions β with a clear grasp of how they enable downstream reporting.
β’ Workflow Management: Modify data processing jobs and workflows as necessary to accommodate evolving business needs.
β’ Code Review & Quality: Evaluate pull requests from other developers, ensuring code quality, performance, and architectural consistency across the codebase.
β’ Environment & Deployment Management: Deploy and promote changes across environments (Dev, QA, UAT, PROD), keeping deployment tracking current, and support environment operations such as restoring environments or tables from another environment or from a specific point in time.
β’ Production Monitoring & Troubleshooting: Monitor and troubleshoot daily production jobs, investigating failures and performance issues using platform-native diagnostic tools, job logs, and table history.
β’ Testing & Automation: Maintain and enhance the automated testing pipeline, including CI/CD workflows and the underlying test framework.
β’ Documentation: Keep technical documentation updated to ensure that institutional knowledge is preserved as the pipelines and connections relied upon by the team evolve.
β’ Technical Reference & Mentoring: Serve as the primary technical reference for the team on the data platform β the person others consult when they require deep platform expertise β and assist fellow team members with data modeling and development topics.
β’ Stakeholder Collaboration: Suggest and recommend architectural and process enhancements, working closely with the client's business and technical stakeholders β including the client's data architecture function β to translate requirements into scalable, well-tested data pipelines, while remaining equally comfortable taking direction from client-side technical leadership.
β’ Extensive experience in data development, with proven hands-on production experience on the Databricks platform.
β’ Strong proficiency in PySpark (DataFrame API, Spark SQL, UDFs, window functions) and Databricks SQL (ANSI SQL, MERGE INTO, COPY INTO, CTEs), including performance tuning techniques such as partition pruning, file compaction, skew handling, and query optimization.
β’ Solid, practical experience with Delta Lake: MERGE/upsert patterns, ACID transactions, time travel, and table maintenance (OPTIMIZE, VACUUM, ZORDER, liquid clustering, Change Data Feed).
β’ Demonstrated experience implementing Slowly Changing Dimensions (Type 1 and Type 2) and dimensional modeling principles (star schema, fact/dimension design) β not requiring you to have designed a model from scratch, but requiring the mindset to understand and extend one.
β’ Experience with medallion (or similar layered) architecture in a production data platform, along with Unity Catalog, jobs/workflows, secrets management, and notebook-based development.
β’ Experience with Git and Azure DevOps (or equivalent) for version control, pull requests, and CI/CD pipelines, in addition to familiarity with Microsoft Azure services (Key Vault, Service Principal/Managed Identity, Data Lake Storage).
β’ Ability to read and navigate a large, established codebase (400+ notebooks), learning and adhering to existing conventions rather than rewriting them, and to quickly adapt in a business-rule-heavy environment.
β’ Advanced English (C1 or above) communication skills, with the capability to work directly with US-based client stakeholders, propose technical recommendations, and align with decisions made by client-side technical leadership.
β’ Nice to Have
β’ Experience with Databricks Asset Bundles or other Infrastructure-as-Code methods for managing jobs, clusters, and permissions as code.
β’ Familiarity with Delta Live Tables and Databricks Genie (AI/BI Genie) for natural-language querying and conversational analytics.
β’ Knowledge of data quality frameworks (e.g., Great Expectations, Soda Core, or custom validation frameworks).
β’ Experience with pytest and databricks-connect for automated testing of Spark pipelines outside of manual notebook execution.
β’ Familiarity with Pydantic or similar typed-configuration methods, and experience with schema migration/versioning strategies (e.g., Flyway, Liquibase, or custom frameworks).
β’ Comfort in using AI-assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to expedite coding, debugging, and code review processes.
β’ Health and dental insurance.
β’ Meal and food allowance.
β’ Childcare assistance.
β’ Extended paternity leave.
β’ Partnerships with gyms and health and wellness professionals via Wellhub (Gympass) TotalPass.
β’ Profit Sharing and Results Participation (PLR).
β’ Life insurance.
β’ Continuous learning platform (CI&T University).
β’ Discount club.
β’ Free online platform dedicated to physical, mental, and overall well-being.
β’ Pregnancy and responsible parenting course.
β’ Partnerships with online learning platforms.
β’ Language learning platform.
β’ And many more!
Commonland
FCamara Consulting & Training
Corndel
Mitsubishi Electric Power Products, Inc.
Get handpicked remote jobs straight to your inbox weekly.