
Principal Data Engineer – RWE
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in India.
• Design and implement data processes for the automated ongoing creation of patient-level data products for dashboards, reports, and research studies.
• Process datasets obtained from data vendors and partners, transforming them into usable formats for Real World Evidence (RWE) studies, dashboards, and other outputs.
• Convert diverse healthcare datasets into reusable data models suitable for observational research and epidemiological studies.
• Transform custom datasets into OMOP format as needed, while addressing any residual data that cannot be converted.
• Construct FAIR-compliant data pipelines and frameworks for semantic data engineering.
• Develop AI-ready datasets that facilitate generative AI applications.
• Collaborate with epidemiologists, statisticians, market access professionals, and health economists to identify requirements and convert them into actionable data structures.
• Work with the RWE programming team to establish data structures and support study outputs.
• Coordinate with IT to ensure that incoming datasets meet the specified requirements.
• Engage with technical personnel from analysis software vendors like Databricks.
• Maintain comprehensive documentation of data flows, schemas, pipelines, and processes.
• Design and execute data validation and monitoring to guarantee accuracy and reliability.
• Resolve issues related to data loading, extraction, and transformation.
• Collaborate with three other members of the Data Engineering team, providing support for workloads as necessary.
• Profound understanding of Real World Data (RWD) and Real World Evidence (RWE) concepts.
• Capability to evaluate business requirements and suggest suitable real-world healthcare datasets for analytical applications.
• In-depth knowledge of healthcare data models and ecosystems.
• Strong expertise in OMOP CDM versions 5.4 and 6, including extensions.
• Familiarity with SNOMED CT, RxNorm, ICD-10, LOINC, and HCPCS/CPT.
• Significant experience in constructing scalable ETL/ELT pipelines.
• Proficiency in Databricks, PySpark, Spark SQL, SQL, and Delta Lake.
• Experience handling large-scale healthcare and patient-level datasets.
• Strong grasp of Semantic Data Engineering principles.
• Experience creating FAIR-compliant data pipelines.
• Familiarity with cloud-based data platforms and distributed processing frameworks.
• Solid Power BI development and data modeling capabilities.
• Ability to develop reusable analytical datasets for dashboards and research studies.
• Experience in designing AI-ready datasets and analytics data products.
• Experience in implementing automated data quality frameworks.
• Strong skills in data profiling, validation, and monitoring.
• Knowledge of healthcare data quality assessment methodologies.
• Excellent stakeholder management and communication abilities.
• Proficient in translating complex business requirements into technical solutions.
• Experience working with cross-functional global teams.
• Exposure to one or more therapeutic areas including Oncology, Respiratory, Immunology & Inflammation, and Infectious Diseases.
• Equal opportunities employer.
• Inclusive and diverse working environment.
• B Corp accredited company committed to social and environmental performance, transparency, and accountability.
• AI-assisted recruitment featuring human-led assessments, selection decisions, and hiring outcomes.
CuraLinc Healthcare
VSP Vision Care
Adoreal
Get handpicked remote jobs straight to your inbox weekly.