
Senior Scientific Data Engineer, R&D Data Platform
Posted Aug 26

Posted Aug 26
This is a fully remote position, open to applicants in Illinois.
• Lead the design and implementation of reusable tools and services for the ingestion, validation, transformation, documentation, discovery, and sharing of scientific data.
• Take ownership of platform capability areas from start to finish, encompassing design, implementation, adoption, operational support, and long-term maintenance.
• Create solutions using Python and SQL, including software packages, data pipelines, APIs, notebooks, workflow utilities, and lightweight internal applications.
• Develop self-service workflows that empower researchers to prepare and share data consistently.
• Collaborate with scientific teams to comprehend studies, analytical workflows, data sources, and technical challenges, converting their needs into a prioritized technical roadmap.
• Establish standards and reusable patterns for organizing and harmonizing data from various sources and promote their adoption.
• Design automated frameworks for data quality and validation.
• Enhance the documentation, traceability, and discoverability of scientific datasets.
• Assess AWS services and features while collaborating with R&D DevOps on architecture, deployment patterns, and operational ownership.
• Develop solutions utilizing Amazon S3, Athena, Glue, EMR, Lambda, and SageMaker.
• Prototype solutions for research programs and generalize successful methods into reusable platform capabilities.
• Provide technical leadership on designs across projects and teams, including conducting design reviews and documenting trade-offs.
• Mentor engineers through code reviews, collaborative work, design feedback, and documentation.
• Assist in the hands-on preparation and analysis of scientific data.
• Apply quantitative and scientific judgment to data and technical solutions.
• Utilize Spark or PySpark for suitable distributed processing tasks.
• Implement software engineering practices including version control, testing, code reviews, documentation, dependency management, continuous integration, and reproducible development.
• Effectively communicate technical concepts, decisions, trade-offs, limitations, and project status to technical, scientific, and leadership audiences.
• Operate autonomously in a dynamic environment by defining ambiguous problems, sequencing tasks, and making justifiable decisions.
• Assist scientists in reducing the manual effort involved in locating, cleaning, interpreting, and restructuring data.
• Enable research teams to employ documented and consistent tools for data preparation and sharing.
• Bachelor's degree in computer science, data science, engineering, statistics, mathematics, bioinformatics, computational science, or another relevant quantitative field.
• Five or more years of relevant professional or applied research experience, or three or more years with an advanced degree in a related area.
• Advanced programming proficiency in Python.
• Strong SQL skills and experience with structured and semi-structured data.
• Proven history of building reusable, maintainable software relied upon by others.
• Experience in designing and delivering several of the following: data pipelines, Python packages, APIs, analytical workflows, notebooks, or internal software tools.
• Extensive hands-on experience utilizing AWS for data processing, analytics, scientific computing, or software development.
• Adequate knowledge of AWS services and architecture to assess technical options, justify design choices, and outline infrastructure requirements.
• Experience conducting or supporting quantitative research, including statistical analysis, machine learning, computational modeling, or other data-intensive research activities.
• Experience in cleaning, integrating, standardizing, or validating data from multiple sources on a significant scale.
• Proficiency with Git, automated testing, technical documentation, code review, and continuous integration.
• Demonstrated capability to investigate ambiguous problems, define an approach, and deliver a workable solution with minimal guidance.
• Experience mentoring or providing technical guidance to fellow engineers, scientists, or analysts.
• Strong communication and collaboration abilities across scientific and technical domains.
• Preferred: Advanced degree in a quantitative, computational, or life-science discipline.
• Preferred: Experience with biomedical, genomic, clinical, proteomic, imaging, laboratory, or other complex scientific data.
• Preferred: Experience supporting research in life sciences, healthcare, diagnostics, or similarly data-intensive and regulated scientific settings.
• Preferred: Production experience with Spark or PySpark and distributed data processing.
• Preferred: Expertise in AWS services such as Athena, Glue, EMR, SageMaker, Lambda, Step Functions, Lake Formation, or related data and analytics technologies.
• Preferred: Experience developing REST APIs or lightweight web applications for non-engineering users.
• Preferred: Experience in designing automated validation frameworks, data contracts, reusable data-processing libraries, or researcher-oriented workflow tools.
• Preferred: Familiarity with metadata management, data cataloging, or data discovery platforms.
• Preferred: Experience with containerization, continuous integration and deployment, or infrastructure as code.
• Preferred: Experience supporting machine-learning workflows or preparing data for model development and evaluation.
• Preferred: Experience working with large files or multimodal datasets.
• Preferred: Awareness of governance considerations for research data.
• Preferred: Experience working within a data mesh, data product, or federated data ownership model.
• Remote work arrangement.
• Standard work shift.
• Travel opportunity: 10% of the time.
Mirantis
Get handpicked remote jobs straight to your inbox weekly.