
Software Engineer – Agentic Data Pipelines
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Design, develop, and maintain agent-based systems that transform biomedical data-source references into reviewed and version-controlled datasets.
• Create LLM-driven pipelines for data sanitization, normalization, and formatting across molecular, genomic, clinical, and literature datasets.
• Establish automated quality-control processes to identify anomalies, highlight inconsistencies, and uphold data standards.
• Assess and refine agent architectures, prompting methodologies, tool specifications, validation cycles, and evaluation frameworks.
• Partner with ML scientists on the Enchant team to convert data needs into scalable acquisition and processing systems.
• Oversee and sustain distributed data pipelines in a production environment, troubleshoot failures, and enhance system resilience.
• Record data provenance, processing choices, and quality metrics to ensure reproducibility and facilitate auditing.
• Safely operate agents utilizing sandboxed execution, least-privilege credentials, restricted network access, and audit logs.
• Identify and communicate potential security vulnerabilities to the team.
• Contribute to the development of a long-term natural-language orchestrator for drug discovery inference, fine-tuning, virtual screenings, and dataset evaluation.
• Master’s degree in a computational STEM discipline, or a Bachelor’s degree with 2+ years of relevant industry experience.
• Proficient Python programming skills and a background in developing and supporting production-level software.
• Practical experience with LLM APIs such as Claude and GPT.
• Familiarity with agentic methodologies, encompassing tool utilization, orchestration, and multi-step reasoning.
• Understanding of biomedical or chemical data sources and formats, including PDB, UniProt, ChEMBL, SDF/MOL, and FASTA.
• Foundational knowledge of data engineering, including ETL design, data validation, and management of structured and unstructured data at scale.
• Experience with Python testing frameworks, particularly pytest fixtures and parameterization.
• Knowledge of agent orchestration frameworks and evaluation systems for LLM-generated code.
• Familiarity with cloud infrastructure and workflow orchestration tools such as AWS, Docker, and Kubernetes.
• Understanding of multimodal biomedical data, including small molecules, proteins, assays, images, ‘omics, and/or clinical records.
• Experience in constructing or curating large-scale datasets for ML model training.
• Awareness of agent security practices, including sandboxing, scoped credentials, and prompt injection techniques.
• Authorized to work in the United States; visa sponsorship details should be included in the application.
• Company-sponsored healthcare.
• Flexible spending accounts.
• Optional life insurance.
• 401K matching.
• Unlimited vacation policy.
• Onsite gym facilities.
• Dining options available.
• New state-of-the-art facility.
• Convenient access to excellent living and recreational areas.
Cypher Consulting Europe S.L.
Sigma Software Group
Coinbase
Get handpicked remote jobs straight to your inbox weekly.