
Data Scientist
Posted 5 hours ago

Posted 5 hours ago
This is a fully remote position, open to applicants in Peru.
• Develop and assess machine learning methodologies for the matching of companies or entities.
• Create embedding and LLM-based matching techniques.
• Formulate scoring and ranking strategies to accurately identify genuine matches while differentiating from duplicates, lookalikes, and unrelated entities.
• Handle complex datasets, including names, aliases, domains, websites, firmographic characteristics, multilingual records, and hierarchical data.
• Establish benchmark datasets, metrics, baselines, and processes for error analysis.
• Design and implement experiments to test hypotheses.
• Evaluate LLM-assisted methodologies against more cost-effective alternatives.
• Investigate model behavior, edge cases, and associated trade-offs.
• Take into account inference economics and scalability from the outset.
• Present experimental results and recommendations to engineering and business stakeholders.
• Independently create experimental pipelines and research methodologies.
• Thoroughly document both successful and unsuccessful experimental outcomes.
• Over 5 years of professional experience in Data Science and Machine Learning.
• Strong foundation in applied Machine Learning principles.
• Proficient in Python and SQL.
• Practical experience with embeddings and semantic similarity techniques.
• Hands-on experience applying LLMs to practical challenges.
• Familiarity with both supervised and unsupervised learning methods.
• Extensive experience in classification and natural language processing (NLP).
• Working knowledge of neural networks and transformer architectures.
• Experience with TensorFlow, PyTorch, PyCaret, or similar ML frameworks.
• Background in retraining or maintaining classification models in a production environment.
• Strong skills in experimental design and model evaluation.
• Experience in defining baselines, metrics, test sets, and conducting error-analysis.
• Capability to assess model quality and demonstrate measurable improvements.
• Solid understanding of scalability and the costs associated with ML inference.
• Proficient in English communication.
• Experience in entity resolution, record linkage, or deduplication (preferred).
• Familiarity with ranking and similarity scoring techniques (preferred).
• Experience with retrieval, clustering, or candidate-generation methods (preferred).
• Knowledge of LLM/embedding solutions optimized for cost and scalability (preferred).
• Familiarity with Spark, Snowflake, Databricks, or BigQuery (preferred).
• Experience dealing with company, domain, website, or firmographic data (preferred).
• Experience working with multilingual datasets (preferred).
• Strong analytical and experimental approach.
• Commitment to intellectual honesty and openness about negative results.
• High degree of autonomy and self-direction.
• Excellent written and verbal communication skills.
• Ability to justify technical recommendations to stakeholders.
• Strong problem-solving abilities.
• Comfort in managing ambiguity and large-scale datasets.
• Ability to balance model quality, cost, and scalability.
• Option for remote work.
• Opportunities for rapid learning and ownership of projects.
• Collaboration with high-performing teams.
• Investment in innovative working methods.
Tendios
GuidePoint Security
Seneca Holdings
Evolent
Get handpicked remote jobs straight to your inbox weekly.