
Data Scientist
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in Peru.
• Develop and assess machine learning methodologies for matching companies/entities.
• Create matching techniques based on embeddings and large language models (LLM).
• Formulate scoring and ranking systems to identify genuine matches while differentiating them from duplicates, lookalikes, and unrelated entities.
• Handle complex datasets, which include names, aliases, domains, websites, firmographic details, multilingual records, and data hierarchies.
• Establish benchmark datasets, metrics, baselines, and processes for error analysis.
• Design and carry out experiments to test hypotheses.
• Evaluate LLM-assisted methods against more cost-effective alternatives.
• Analyze model performance, edge cases, and trade-offs.
• Consider the economics of inference and scalability from the outset.
• Share experimental results and recommendations with engineering and business stakeholders.
• Independently create experimental workflows and research strategies.
• Thoroughly document both successful and unsuccessful experiments.
• Over 5 years of professional experience in Data Science / Machine Learning.
• Solid foundation in applied Machine Learning principles.
• Proficient in Python and SQL.
• Practical experience with embeddings and semantic similarity.
• Hands-on application of LLMs to real-world challenges.
• Experience with both supervised and unsupervised learning techniques.
• Strong background in classification and natural language processing (NLP).
• Familiarity with neural networks and transformer architectures.
• Proficient in using TensorFlow, PyTorch, PyCaret, or similar ML frameworks.
• Experience in retraining or managing classification models in a production environment.
• Strong skills in experimental design and model evaluation.
• Experience in defining baselines, metrics, test sets, and conducting error analysis.
• Ability to assess model quality and demonstrate quantifiable enhancements.
• Deep understanding of scalability and costs associated with ML inference.
• Excellent communication skills in English.
• Nice-to-have: experience in entity resolution, record linkage, or deduplication.
• Nice-to-have: familiarity with ranking and similarity scoring.
• Nice-to-have: knowledge of retrieval, clustering, or candidate-generation techniques.
• Nice-to-have: LLM/embedding solutions tailored for cost and scalability constraints.
• Nice-to-have: experience with Spark, Snowflake, Databricks, or BigQuery.
• Nice-to-have: background in company, domain, website, or firmographic data.
• Nice-to-have: experience in working with multilingual datasets.
• Strong analytical and experimental mindset.
• Intellectual honesty and a willingness to communicate unfavorable results.
• High level of autonomy and self-direction.
• Excellent written and verbal communication abilities.
• Capability to defend technical recommendations to stakeholders.
• Strong problem-solving skills.
• Comfort with ambiguity and large-scale datasets.
• Ability to balance model quality, cost, and scalability.
• Option for remote work.
InductiveHealth Informatics
Autodesk
Bixal
Mondelēz International
Get handpicked remote jobs straight to your inbox weekly.