
Senior Scientific Data Engineer, R&D Data Platform
Posted Aug 26

Posted Aug 26
This is a fully remote position, open to applicants in Illinois.
• Take charge of the design and implementation of reusable tools and services for the ingestion, validation, transformation, documentation, discovery, and sharing of scientific data.
• Manage platform capability areas comprehensively, covering design, implementation, adoption, operational support, and sustainable maintenance.
• Create solutions using Python and SQL, including software packages, data pipelines, APIs, notebooks, workflow utilities, and lightweight internal applications.
• Develop self-service workflows that enable researchers to prepare and share data systematically.
• Collaborate with scientific teams to comprehend studies, analytical workflows, data sources, and technical challenges, converting needs into a prioritized roadmap.
• Set standards and reusable patterns for organizing and harmonizing diverse data sources.
• Design automated frameworks for data quality and validation.
• Enhance documentation, traceability, and discoverability of scientific datasets.
• Assess AWS services and features, collaborating with R&D DevOps on architecture, deployment, and operational responsibilities.
• Build solutions utilizing Amazon S3, Athena, Glue, EMR, Lambda, and SageMaker.
• Prototype solutions for research programs and generalize successful strategies into reusable platform capabilities.
• Offer technical leadership through design reviews, trade-off documentation, and alignment across teams.
• Guide engineers through code reviews, collaborative work, design feedback, and documentation.
• Assist in the hands-on preparation and analysis of scientific data.
• Employ quantitative and scientific judgment in data and technical solutions.
• Utilize Spark or PySpark for handling large or computationally intensive datasets.
• Implement software engineering best practices, including version control, testing, code reviews, documentation, dependency management, continuous integration, and reproducible development.
• Convey technical concepts, decisions, trade-offs, limitations, and project status to technical, scientific, and leadership audiences.
• Work independently in a dynamic environment while keeping stakeholders informed.
• Bachelor’s degree in computer science, data science, engineering, statistics, mathematics, bioinformatics, computational science, or another relevant quantitative field.
• At least five years of relevant professional or applied research experience, or three years with an advanced degree in a pertinent area.
• Advanced proficiency in Python programming.
• Strong SQL skills and experience with both structured and semi-structured data.
• Proven experience in building reusable, maintainable software that is relied upon by others.
• Experience in designing and delivering data pipelines, Python packages, APIs, analytical workflows, notebooks, or internal software tools.
• Significant hands-on experience with AWS for data processing, analytics, scientific computing, or software development.
• Adequate knowledge of AWS services and architecture to assess technical options, justify design choices, and outline infrastructure needs.
• Experience in conducting or supporting quantitative research.
• Familiarity with cleaning, integrating, standardizing, or validating data from various sources at scale.
• Proficiency with Git, automated testing, technical documentation, code reviews, and continuous integration.
• Ability to explore ambiguous issues, devise an approach, and deliver a functional solution with minimal guidance.
• Experience in mentoring or providing technical advice to engineers, scientists, or analysts.
• Excellent communication and collaboration abilities.
• Preferred: advanced degree in a quantitative, computational, or life-science discipline.
• Preferred: experience with biomedical, genomic, clinical, proteomic, imaging, laboratory, or other complex scientific data.
• Preferred: experience in life sciences, healthcare, diagnostics, or a similarly data-intensive and regulated scientific setting.
• Preferred: practical experience with Spark or PySpark.
• Preferred: expertise in AWS services such as Athena, Glue, EMR, SageMaker, Lambda, Step Functions, Lake Formation, or related technologies.
• Preferred: familiarity with REST APIs or lightweight web applications.
• Preferred: experience with automated validation frameworks, data contracts, reusable data-processing libraries, or researcher-facing workflow tools.
• Preferred: knowledge of metadata management, data catalog, or data discovery platforms.
• Preferred: experience with containerization, CI/CD, or infrastructure as code.
• Preferred: knowledge of machine-learning workflows or model-development data preparation.
• Preferred: experience with large files or multimodal datasets.
• Preferred: understanding of research data governance, access control, de-identification, and sensitive clinical information.
• Preferred: familiarity with data mesh, data product, or federated data ownership models.
• Remote work arrangement.
• Travel required: 10% of the time.
Mirantis
Get handpicked remote jobs straight to your inbox weekly.