
Senior Databricks Engineer
Posted Aug 28

Posted Aug 28
This is a fully remote position, open to applicants in United States.
• Design, develop, test, and deploy scalable data pipelines and transformation workflows using Databricks.
• Analyze and reverse engineer existing AWS-based data processing solutions to uncover transformations, business rules, dependencies, orchestration, and integration needs.
• Construct, enhance, and maintain the Bronze, Silver, and Gold data layers.
• Create and implement Databricks Auto Loader ingestion solutions featuring schema management, checkpointing, incremental processing, backfill/reprocessing, and production-scale file ingestion.
• Develop and manage Lakeflow Declarative Pipelines for both batch and streaming workloads, ensuring data quality, quarantine/error handling, dependencies, monitoring, and recovery.
• Build real-time and batch processing solutions utilizing Databricks, Structured Streaming, and associated technologies.
• Apply transformation logic using Apache Spark, PySpark, Spark SQL, and Delta Lake.
• Design and optimize Delta Lake solutions utilizing MERGE, Change Data Feed, partitioning/clustering, retention, and optimization strategies.
• Ensure reliable ingestion, transformation, enrichment, and delivery of high-volume customer data.
• Implement Unity Catalog governance structures and facilitate asset promotion across development, testing, and production environments.
• Develop and maintain integrations among Databricks, Adobe, AWS services, and downstream systems.
• Analyze AWS Glue, S3, Redshift, Redshift Spectrum, and related AWS services to assess current functionality and target Databricks implementations.
• Establish reusable Databricks workflows, deployment-as-code patterns, frameworks, utilities, and engineering standards.
• Adhere to best practices for data quality, performance, scalability, observability, reliability, security, and maintainability.
• Optimize Spark workloads, pipelines, queries, streaming processes, and data structures.
• Diagnose complex pipeline, integration, data quality, performance, and production issues; engage in root-cause analysis and remediation efforts.
• Create automated validation and testing methodologies for data products.
• Execute legacy-to-target parity validation and resolve discrepancies.
• Assess legacy logic to retain business functionality while eliminating unnecessary technical limitations.
• Participate in design discussions, peer code reviews, technical evaluations, and solution refinement.
• Comply with client development, CI/CD, security, governance, and deployment protocols.
• Generate and maintain technical documentation.
• Provide practical knowledge transfer and mentoring to client engineers.
• Utilize approved AI-assisted development tools as appropriate.
• Identify and suggest automation and engineering enhancements.
• A minimum of 6 years of experience in data engineering or software engineering, with substantial expertise in designing, developing, and supporting enterprise-scale data platforms.
• At least 3 years of practical experience with Databricks, focusing on the development and operation of production data engineering solutions.
• Proven production experience with Databricks Auto Loader, encompassing incremental cloud object storage ingestion, schema inference and evolution, schema hints, rescued data handling, checkpoint/state management, backfill and reprocessing strategies, and scalable file discovery.
• Demonstrated production experience with Lakeflow Declarative Pipelines (formerly Delta Live Tables), including batch and streaming pipelines, data quality standards, failed record/quarantine management, streaming tables and materialized views, refresh processing, dependency management, monitoring, and alerting.
• Hands-on experience implementing Unity Catalog, covering catalog/schema/volume design, external locations, storage credentials, grants, row- and column-level access controls, lineage, and asset promotion.
• Practical experience with Delta Lake, including MERGE patterns, Change Data Feed, time travel, OPTIMIZE, Z ORDER and/or liquid clustering, VACUUM and retention policies, and partitioning strategies.
• Proficient in Apache Spark, PySpark, and Spark SQL, with experience in complex transformations and production performance tuning.
• Ability to diagnose and enhance Spark workloads through partition sizing, shuffle optimization, join strategy, skew handling, and Spark UI analysis.
• Experience with production Structured Streaming, including watermarking, handling late-arriving data, stateful processing, checkpointing, and considerations for exactly-once processing.
• Knowledge of designing and implementing Medallion Architecture, with personal experience in building Bronze, Silver, and Gold data layers.
• Familiarity with Databricks Workflows and deployment as code, including managing dependencies, retry/failure handling, alerting, and Git-based environment promotion using Databricks Asset Bundles, Terraform, or similar automation tools.
• Strong understanding of distributed data processing, data optimization, scalability, and resilience of production pipelines.
• Experience with production-level data quality, error handling, monitoring, logging, observability, and pipeline recovery.
• Familiarity with CI/CD, automated deployment, Git-based source control, branching, pull requests, and peer code reviews.
• Strong hands-on experience with AWS Glue ETL using PySpark and Python.
• Ability to analyze unfamiliar and undocumented AWS Glue code to identify transformations, data movement, dependencies, and business rules.
• Solid knowledge of Amazon S3.
• Experience with Amazon Redshift, including data structures, distribution and sort strategies, COPY/UNLOAD patterns, and stored procedures.
• Strong understanding of AWS IAM and data access patterns.
• Working knowledge of AWS Glue Crawlers and Glue Data Catalog.
• Familiarity with Redshift Spectrum and external S3-backed data access patterns.
• Knowledge of Lambda, Step Functions, EventBridge, Athena, CloudWatch, and Secrets Manager.
• Ability to trace end-to-end AWS data pipelines across multiple services.
• Experience in reverse engineering undocumented legacy data pipelines.
• Ability to identify pipeline behavior and dependencies without original developers or complete documentation.
• Experience with data reconciliation and parity validation between legacy and rebuilt pipelines.
• Capability to differentiate business logic from legacy technical workarounds.
• Ability to troubleshoot intricate data engineering and production issues independently.
• Experience in Agile delivery environments, successfully meeting prioritized product backlogs.
• Excellent communication and collaboration skills across engineering, architecture, product, and business teams.
• Ability to successfully pass a background investigation for U.S. employment.
• Competitive salary with a profit participation program.
• Comprehensive medical, dental, and vision coverage.
• Basic life insurance and accidental death & dismemberment coverage.
• Matching contributions available through a 401(k) plan.
• CGI share purchase plan.
• Paid accrued vacation leave ranging from 10 to 20 days annually.
• 10 paid holidays each year.
• At least 80 consecutive hours of paid sick/safe leave.
• Paid parental leave ranging from 20 to 70 consecutive business days.
• Bereavement leave ranging from 1 to 7 days each year.
• Paid jury duty leave up to the time summoned.
• Opportunities for learning and tuition assistance.
• Wellness and well-being programs.
University Hospitals Urgent Care
WellStreet Urgent Care
Prisma Health Urgent Care
Canadian Solar Inc.
Get handpicked remote jobs straight to your inbox weekly.