
Data Engineer, COA Accelerator
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Spain, +5 more countries.
• Design, construct, and sustain data infrastructure that underpins IQVIA’s AI-driven COA strategy and COA Accelerator functionalities.
• Take responsibility for the ingestion, transformation, normalization, enrichment, indexing, versioning, and governance of public, proprietary, and client-specific data sources.
• Create ingestion pipelines for both structured and unstructured sources, including documents, databases, APIs, clinical trial registries, regulatory documents, scientific publications, and internal repositories.
• Convert source material into standardized, searchable, AI-ready formats for evidence retrieval, source citation, recommendation generation, and expert review.
• Establish processes for document parsing, OCR, text extraction, metadata enrichment, chunking, deduplication, versioning, indexing, and quality control.
• Contribute to the platform’s core knowledge layer, encompassing COA metadata, psychometric evidence, therapeutic area mappings, endpoint usage, regulatory precedents, and scientific evidence.
• Facilitate the integration of public evidence sources, including clinical trial registries, FDA labels, EMA EPARs, HTA records, scientific literature, FDA guidance, and qualification documents.
• Assist with the ingestion of proprietary internal knowledge, publications, thought leadership, and expert-authored content.
• Prepare data for retrieval-augmented generation utilizing chunking, embeddings, indexes, metadata filters, and source reference structures.
• Collaborate with AI engineers to enhance retrieval precision, recall, relevance, and citation accuracy.
• Implement hybrid retrieval methods that combine semantic search, keyword search, structured database queries, and metadata filtering.
• Ensure traceability between AI-generated outputs and source documents.
• Establish data quality controls and maintain audit trails for ingestion, transformation, updates, deletions, access rights, and downstream usage.
• Collaborate with legal, security, compliance, product, and domain stakeholders on contractual, licensing, privacy, intellectual property, and governance needs.
• Work in partnership with AI Engineering, Product, COA Science, and Software Engineering teams.
• Bachelor’s degree in computer science, data engineering, data science, information systems, bioinformatics, computational biology, engineering, or a related technical discipline.
• Proven experience in designing, building, and maintaining data pipelines for both structured and unstructured data.
• Proficient in Python and SQL.
• Familiarity with APIs, relational databases, document stores, search indexes, cloud data platforms, and ETL/ELT workflows.
• Experience managing large volumes of text-heavy documents.
• Strong understanding of data cleaning, normalization, metadata management, document parsing, indexing, lineage, versioning, and auditability.
• Knowledge of data modeling for complex scientific, clinical, regulatory, or healthcare knowledge domains.
• Capability to translate domain expert requirements into practical data structures, metadata models, retrieval-ready content, and maintainable pipelines.
• Excellent attention to detail and ability to identify data quality issues.
• Ability to collaborate effectively with AI engineers, product managers, COA scientists, software engineers, security stakeholders, legal teams, and commercial teams.
• Strong documentation skills.
• Experience in life sciences, clinical research, healthcare, regulatory data, scientific publishing, HEOR, clinical outcome assessments, patient-reported outcomes, or medical evidence management is highly preferred.
• Familiarity with clinical trial registries, regulatory labels, HTA reports, scientific literature databases, medical knowledge repositories, or similar evidence sources.
• Experience with vector databases, embeddings, semantic search, Elasticsearch/OpenSearch, Azure AI Search, Pinecone, Weaviate, Milvus, Qdrant, or similar technologies.
• Proficiency in cloud data platforms such as Azure, AWS, or GCP.
• Experience with document AI, OCR, layout-aware parsing, table extraction, metadata enrichment, taxonomy development, or controlled vocabularies is a plus.
• Understanding of ontology development, biomedical terminologies, controlled vocabularies, evidence classification, and structured knowledge representation.
• Knowledge of GDPR, data privacy, intellectual property constraints, licensed content, confidential client data, access control, and secure data handling.
• Ability to work independently in a remote or hybrid environment while collaborating with global teams.
• Significant experience utilizing AI tools in professional settings.
• Fluent in English.
• A rewarding and progressive career path.
• Comprehensive training and support.
• Opportunities for development and growth.
• Hands-on influence in creating and delivering innovative solutions.
• A multi-cultural, collegial, and collaborative work environment.
SysMap Solutions
Get handpicked remote jobs straight to your inbox weekly.