
Clinical Data Scientist
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in United States.
β’ Configure and uphold EDC forms and eCRF specifications for data ingestion and subsequent analysis.
β’ Oversee comprehensive data management for Evidence Generation studies from ingestion to cleaning, quality control, and data lock.
β’ Develop, validate, and manage research databases.
β’ Create and execute data-cleaning workflows, which include queries, reconciliation, and audit trails.
β’ Extract, parse, and analyze EHR data from FHIR bundles, HL7v2 messages, and flat-file/CSV exports.
β’ Construct Python/SQL transformation pipelines that clean, standardize, and reshape raw EHR extracts into validated EDC-ready datasets.
β’ Implement field mapping, unit harmonization, deduplication, and derivation logic.
β’ Maintain reproducible, version-controlled transformation code.
β’ Design and manage data models for automated EHR-to-EDC ingestion.
β’ Collaborate with APIs/integration endpoints and harmonize diverse data models.
β’ Collaborate with site IT/informatics and Product teams regarding EHR constraints, extract structures, and change management.
β’ Translate EHR data realities into study data capture and monitoring plans.
β’ Develop and maintain Data Management Plans (DMPs), Statistical Analysis Plans (SAPs), Standard Operating Procedures (SOPs), runbooks, transformation/mapping specifications, and training materials.
β’ Conduct descriptive, comparative, and time-to-event statistical analyses.
β’ Utilize data-science techniques to derive features/cohorts, perform exploratory analysis, and assess data quality.
β’ Work collaboratively with stakeholders and train team members through SOPs, runbooks, and templates.
β’ Enhance EHR-to-Visualization-to-EDC auto-import pipelines and assist with sponsor-ready reporting.
β’ 5β7+ years of experience in clinical research/real-world evidence (RWE) data management, including hands-on database building and cleaning/quality control ownership.
β’ Proficiency in SQL and Python (or R) for data transformation, ETL pipeline development, quality control checks, and reproducible pipelines.
β’ Practical experience in transforming raw EHR extracts (FHIR, HL7v2, or flat-file exports) into structured, analysis/EDC-ready datasets.
β’ Familiarity with EHR or EHR-derived datasets.
β’ Knowledge of ICD-10, CPT, and LOINC; with RxNorm being preferred.
β’ Understanding of API-based ingestion and the integration of multiple data sources/models.
β’ Experience in creating and maintaining DMPs, SOPs, and runbooks.
β’ Proven experience with audit-ready documentation.
β’ Working knowledge of version control and reproducible pipeline practices, such as Git.
β’ Familiarity with the integration of leading AI techniques into work products.
β’ Experience with EDC platforms such as Medidata Rave, REDCap, Castor, or Veeva is preferred.
β’ Familiarity with CDISC (SDTM/ADaM) or equivalent standardization experience is preferred.
β’ Statistical experience in real-world/implementation studies is preferred.
β’ Experience with pipeline orchestration/data engineering tools and AWS data warehouses is preferred.
β’ Must be able to work in the United States without needing sponsorship, either now or in the future.
β’ Equity
β’ Performance-based bonus
β’ Medical insurance
β’ Dental insurance
β’ Vision insurance
β’ 401(k)
β’ Generous vacation
β’ Additional benefits for full-time employees
β’ Flexible remote work arrangement
Lightcast
Allstate
ACT
Working Families Party
Get handpicked remote jobs straight to your inbox weekly.