
Senior Data Scientist, AI Retrieval Systems
Posted Aug 26

Posted Aug 26
This is a fully remote position, open to applicants in United States.
• Model biomedical knowledge for research on rare diseases by integrating disease and phenotype ontologies as well as controlled vocabularies into PostgreSQL.
• Maintain release and refresh pathways, reconcile identifiers across various sources, and assess term hierarchies for clinical significance.
• Develop retrieval-augmented services that connect everyday language with clinical concepts.
• Optimize keyword and vector searches across extensive biomedical corpora while managing recall and latency trade-offs.
• Create ranking and relevance layers that incorporate domain-specific weighting and ensure graceful degradation.
• Deliver user interfaces using Next.js, React, and TypeScript, which include question and confirmation flows, result presentations, and live pipeline statuses.
• Continuously deploy to NIH on-premises and high-performance computing Kubernetes environments utilizing Helm charts, StatefulSets, secrets, ingress, GPU scheduling, and scheduled jobs.
• Construct and sustain evaluation systems as regression suites for retrieval and concept mapping.
• Capture request identifiers, latency, errors, and selected concepts for the review of AI-assisted results.
• Translate the needs of researchers, clinicians, and patient communities into data models, retrieval behavior, and interface design.
• Contribute to manuscripts, conference abstracts, and posters alongside NIH investigators, receiving authorship credit.
• Collaborate directly with NIH program staff, clinical geneticists, and specialists in rare disease information.
• Bachelor’s degree in Data Science, Computer Science, Bioinformatics, Biomedical Informatics, or a related discipline; equivalent professional experience may be accepted in lieu of a degree.
• Advanced degree is preferred.
• Minimum of 5 years of experience in building and operating production software or data systems.
• At least 2 years of experience in delivering LLM-powered applications relied upon by users.
• Experience in developing retrieval systems from start to finish, including indexing, query formulation, and retrieval-quality assessment against actual data.
• Familiarity with evaluating systems that do not have a singular correct answer using golden sets, offline regression suites, Recall@K, or MRR.
• Experience with structured output and tool/function invocation using typed schemas and validation.
• Capability to manage a service comprehensively, from schema design through deployment and operation.
• Ability to obtain and maintain a Public Trust Security clearance.
• Proficiency in Python with FastAPI, Pydantic, and pytest.
• Experience with PostgreSQL, including vector search, full-text search, embedding pipelines, indexing, and query optimization.
• Knowledge in LLM application engineering, including provider APIs and gateways, prompt/context design, structured generation, and tool utilization.
• Experience with data ingestion and transformation pipelines that have a repeatable refresh process.
• Familiarity with containers and Kubernetes.
• Proficiency in Next.js, React, and TypeScript.
• Experience with Git-based collaboration and CI/CD practices.
• 100% Medical, Dental & Vision Coverage for Employees.
• Paid Time Off and Paid Holidays.
• 401K match up to 5%.
• Educational Benefits for Career Growth.
• Employee Referral Bonus.
• Flexible Spending Accounts: Healthcare (FSA), Parking Reimbursement Account (PRK), Dependent Care Assistant Program (DCAP), Transportation Reimbursement Account (TRN).
Sigma Software Group
Sigma Software Group
BIP Brasil
Get handpicked remote jobs straight to your inbox weekly.