Remotery

Senior Research Data Engineer

atPointClickCareRemoteCA flagCanadaFull-timeData EngineerSeniorC$159.1k – C$176.7k/year

Posted 4 days ago

This is a fully remote position, open to applicants in Canada.

📋 Description

• Take ownership of the gold data layer by transforming disorganized silver tables into well-curated, semantically rich, clean, and documented gold datasets that are suitable for AI model development, including datasets and features that can be reused across various AI projects.

• Ensure that the data evolves as products and needs change. This involves:

• Reverse-engineering data semantics. Collaborate with product engineers, clinical experts, and workflow specialists to understand how the products are used and how data is generated in real-world scenarios.

• Comprehending SQL queries, stored procedures, technical data definitions, and other coding elements to grasp how products represent and modify data.

• Gaining insights into how data is ingested into the data lake, what silver tables and columns represent, and their behavior.

• Documenting provenance, semantics, clinical event sequencing, cross-module record linkage, and known anomalies.

• Aligning semantics with AI requirements. Understand researchers' data needs to design and develop the gold data product, including evolving documentation that meets AI applied research demands, creating an efficient AI-first foundation for model R&D.

• Curating datasets across various modalities. For different AI applications such as generative AI, RAG, predictive techniques, and others, support researchers' requirements for chunked and tagged unstructured content with rich metadata, point-in-time-correct features, and clean labels. For classical ML and statistical analysis, deliver model-ready tables.

• Developing reusable pipelines. Create transformations from silver to gold within Databricks/Spark as scheduled, observable workloads. Design these so that researchers can iterate on new features and data mixes without the need to rebuild from scratch.

• Automating quality control, filtering, and synthesis. Address research needs for programmatic labeling, weak supervision, near-duplicate detection, boilerplate and noise removal, and LLM-API-driven synthetic data generation where ground truth is limited.

• Versioning and handoff. Maintain reproducible dataset snapshots and define clear lineage and semantic definitions so that the downstream team can effectively use and reuse gold datasets for AI R&D.


⛳️ Requirements

• Over 5 years of experience in building production data systems, with at least 2 years supporting ML or AI workloads.

• Proven ability to quickly learn complex new data domains through reading source code, interviewing experts, and creating durable artifacts that others depend on.

• Proficiency in Python, SQL, and PySpark/Databricks for managing large, complex datasets.

• Expertise in SQL, particularly in reading complex stored procedures and reverse-engineering business logic from queries.

• In-depth knowledge of the Databricks ecosystem: Delta Lake, Unity Catalog, Spark/PySpark tuning, and MLflow.

• Familiarity with AI concepts: a working understanding of embeddings, tokenization, feature engineering, point-in-time correctness, train/validation/test splits, data drift, and the distinctions between classical ML and generative model data requirements.

• Experience in data wrangling across modalities: transforming unstructured content (such as text, PDFs, transcripts, logs) and structured tabular data into clean, model-ready formats.

• Knowledge of AI-friendly data formats (e.g., Parquet, Hugging Face datasets) and storage layout decisions—including partitioning, sharding, and caching—that ensure researcher workflows are responsive in Azure, AWS, or other environments.

• Expertise in data quality, filtering, and synthesis pipelines: supporting programmatic labeling and weak supervision (e.g., Snorkel or equivalent), near-duplicate detection (MinHash/LSH), content and quality filters, and LLM-API-driven synthetic data generation.

• Experience with pipeline orchestration tools (such as Airflow, Databricks Workflows, Dagster, or Prefect) and dataset versioning, including Unity Catalog and feature-store support.

• Background in handling regulated or sensitive data under controlled access (e.g., HIPAA or equivalent) and familiarity with general de-identification concepts.

• Proficient in Git-based version control and CI/CD for data and code.

• Strong written documentation skills, with the ability to elicit requirements and tacit knowledge from both technical and non-technical experts.

• A Bachelor’s degree in computer science, data science, engineering, statistics, or a related field. Equivalent practical experience will also be considered.


🏝️ Benefits

• Benefits starting from Day 1!

• Retirement Plan Matching

• Flexible Paid Time Off

• Wellness Support Programs and Resources

• Parental & Caregiver Leaves

• Fertility & Adoption Support

• Continuous Development Support Program

• Employee Assistance Program

• Allyship and Inclusion Communities

• Employee Recognition … and more!

People also viewed

Redhorse Corporation3 hours ago

Senior Full-Stack Software Engineer, Data & Migrations

US flagUnited States OnlyFull-timeData Engineer
ApplyView job
GE Aerospace3 hours ago

Senior AI Data Engineer

US flagUnited States OnlyFull-timeData Engineer$95k – $159k/year
ApplyView job
Sayari3 hours ago

Senior Data Engineer

US flagUnited States OnlyFull-timeData Engineer$120k – $145k/year
ApplyView job
Minor Hotels Europe and Americas3 hours ago

Data Engineer

MX flagMexico OnlyFull-timeData Engineer
ApplyView job
Samsara3 hours ago

Senior Data Engineer

US flagCalifornia OnlyFull-timeData Engineer$134.5k – $203.4k/year
ApplyView job
Trulieve4 hours ago

Senior Data Engineer

US flagFlorida OnlyFull-timeData Engineer
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers