
Senior Data Developer, Databricks
Posted Jul 31

Posted Jul 31
This is a fully remote position, open to applicants in Colombia.
• Address assigned tickets and bugs while consistently seeking improvement opportunities beyond the immediate task — proactively identifying technical debt, refactoring opportunities, and automation gaps instead of limiting contributions to assigned work.
• Take ownership of and enhance notebooks across the Stage, Bronze, Silver, and Gold layers, from data ingestion to the dimensional model utilized by business intelligence reporting tools.
• Design and develop dimensional modeling artifacts — facts, dimensions, and slowly changing dimensions — with a clear understanding of how they facilitate downstream reporting.
• Modify data processing jobs and workflows as necessary to accommodate evolving business requirements.
• Review pull requests from fellow developers, ensuring code quality, performance, and architectural consistency throughout the codebase.
• Deploy and promote changes across various environments (Dev, QA, UAT, PROD), maintaining up-to-date deployment tracking, and support environmental operations such as restoring environments or tables from another environment or a specific point in time.
• Monitor and troubleshoot daily production jobs, investigating failures and performance issues using platform-native diagnostic tools, job logs, and table history.
• Maintain and enhance the automated testing pipeline, including CI/CD workflows and the foundational test framework.
• Ensure technical documentation remains current to preserve institutional knowledge as the pipelines and connections evolve.
• Serve as the primary technical resource for the team on the data platform — the person others consult for deep platform expertise — and assist team members with data modeling and development topics.
• Suggest and endorse architectural and process enhancements, working with the client's business and technical stakeholders — including the client's data architecture function — to translate requirements into scalable, well-tested data pipelines, while also being comfortable taking guidance from client-side technical leadership.
• Extensive experience in data development, with proven hands-on production expertise on the Databricks platform.
• Strong proficiency in PySpark (DataFrame API, Spark SQL, UDFs, window functions) and Databricks SQL (ANSI SQL, MERGE INTO, COPY INTO, CTEs), including performance tuning techniques such as partition pruning, file compaction, skew handling, and query optimization.
• Solid practical experience with Delta Lake: MERGE/upsert patterns, ACID transactions, time travel, and table maintenance (OPTIMIZE, VACUUM, ZORDER, liquid clustering, Change Data Feed).
• Demonstrated experience implementing Slowly Changing Dimensions (Type 1 and Type 2) and dimensional modeling concepts (star schema, fact/dimension design) — not requiring you to have created a model from scratch, but necessitating the mindset to understand and extend one.
• Experience with medallion (or similar layered) architecture in a production data platform, along with Unity Catalog, jobs/workflows, secrets management, and notebook-based development.
• Familiarity with Git and Azure DevOps (or equivalent) for version control, pull requests, and CI/CD pipelines, as well as Microsoft Azure services (Key Vault, Service Principal/Managed Identity, Data Lake Storage).
• Ability to navigate and comprehend a large, established codebase (400+ notebooks), learning and adhering to existing conventions rather than rewriting them, and to quickly ramp up in a business-rule-heavy environment.
• Advanced English (C1 or above) communication skills, enabling direct collaboration with US-based client stakeholders, proposing technical recommendations, and aligning with decisions made by client-side technical leadership.
• Nice to Have: Experience with Databricks Asset Bundles or other Infrastructure-as-Code methodologies for managing jobs, clusters, and permissions as code; Familiarity with Delta Live Tables and Databricks Genie (AI/BI Genie) for natural-language querying and conversational analytics; Familiarity with data quality frameworks (e.g., Great Expectations, Soda Core, or custom validation frameworks); Experience with pytest and databricks-connect for automated testing of Spark pipelines outside of manual notebook execution; Familiarity with Pydantic or similar typed-configuration approaches, and experience with schema migration/versioning methods (e.g., Flyway, Liquibase, or custom frameworks); Proficient in using AI-assisted development tools (e.g., GitHub Copilot, Cursor, or similar) to enhance coding, debugging, and code review processes.
• Premium Healthcare
• Meal voucher
• Maternity and Parental leaves
• Mobile services subsidy
• Sick pay-Life insurance
• CI&T University
• Colombian Holidays
• Paid Vacations
Social Discovery Group
NatWest Group
Towa Software
Clinical Outcomes Solutions
Get handpicked remote jobs straight to your inbox weekly.