
Data Engineer
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in Canada.
• Design, construct, and maintain ingestion pipelines from high-volume external APIs that operate continuously and reliably at scale.
• Develop ingestion and transformation workflows in Databricks utilizing Spark/PySpark, SQL, and Delta Live Tables.
• Implement medallion architecture patterns to transform raw content into clean, structured, and analysis-ready data.
• Collaborate with the Data Scientist to establish deduplication and relevance-filtering infrastructure.
• Apply schema evolution management and data validation rules.
• Configure and oversee Delta Lake storage structures, including tables, partitions, and optimization routines.
• Design and refine data schemas to balance query performance, cost, and maintainability.
• Maintain documentation for metadata and table structures for Data Science and application teams.
• Ensure pipeline reliability and observability through effective error handling, retries, monitoring, and alerting.
• Adapt pipelines in response to changes in external API contracts, rate limits, authentication methods, and new data sources.
• Diagnose pipeline failures, execute recovery, and optimize performance.
• Create, schedule, and monitor workflows using Databricks Workflows, Delta Live Tables, or similar tools.
• Contribute to CI/CD pipelines for code deployment, versioning, and environment management.
• Work alongside the Data Scientist to deliver structured data for LLM/NLP pipelines and downstream models.
• Engage in data architecture discussions and propose solutions as team needs evolve.
• Document pipelines, data dictionaries, job schedules, and transformation logic.
• Assist in onboarding new data sources and pipelines as the product grows.
• Preference for candidates based in Quebec.
• Proficiency in French (both spoken and written) is a significant asset along with English.
• 3 to 5 years of experience in data engineering, with a strong background in building and managing production-grade data pipelines.
• Familiarity with data modeling, data quality, and schema evolution.
• Strong understanding of data pipeline reliability practices, including monitoring, alerting, and gracefully handling failures in a continuously running system.
• Practical experience with Databricks or a similar Spark-based environment, focusing on schema design, Delta Lake, performance tuning, and pipeline orchestration.
• Experience with at least one major cloud provider; Azure is preferred, while AWS/GCP experience is also beneficial.
• Experience with integrating external APIs at scale, covering aspects like authentication, pagination, rate limiting, retries, and error handling.
• Advanced proficiency in Python and SQL.
• Comfortable working with unstructured and semi-structured text data at scale.
• Familiarity with LLM prompting and/or a basic understanding of AI/NLP concepts.
• Exposure to medallion architecture or lakehouse best practices.
• Experience with orchestration frameworks such as ADF, Workflows, Airflow, or DBX.
• Familiarity with CI/CD tools and version control systems, such as Git or GitHub Actions.
• Basic understanding of security practices, including RBAC, encryption, and credential management.
• Databricks certification (Data Engineer Associate or equivalent) is preferred.
• Competitive compensation package based on experience and qualifications.
• Medical, Dental, and Vision Insurance.
• 401(k) Plan with Company Match.
• Generous Paid Time Off (PTO).
• Company-Paid Holidays.
• Flexible Work Options / Work-from-home opportunities.
• On-Call Compensation — Additional pay for eligible on-call shifts.
Railroad19
GFT Technologies
Get handpicked remote jobs straight to your inbox weekly.