
Data Scientist
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in Peru.
• Develop and assess machine learning strategies for matching companies/entities.
• Create embedding and LLM-based matching techniques.
• Formulate scoring and ranking methods to pinpoint genuine matches while differentiating them from duplicates, lookalikes, and unrelated entities.
• Manage complex data, including names, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies.
• Establish benchmark datasets, metrics, baselines, and error-analysis procedures.
• Design and implement experiments to test hypotheses.
• Compare LLM-assisted methods with more cost-effective alternatives.
• Evaluate model performance, edge cases, and trade-offs.
• Incorporate inference economics and scalability considerations from the outset.
• Present experimental results and recommendations to engineering and business stakeholders.
• Independently create experimental pipelines and research methodologies.
• Thoroughly document both successful and failed experiments.
• Minimum of 5 years of professional experience in Data Science / Machine Learning.
• Solid foundation in applied Machine Learning principles.
• Proficient in Python and SQL.
• Practical experience with embeddings and semantic similarity.
• Hands-on experience applying LLMs to real-world challenges.
• Familiarity with supervised and unsupervised learning techniques.
• Substantial experience in classification and natural language processing (NLP).
• Basic knowledge of neural networks and transformer architectures.
• Practical experience with TensorFlow, PyTorch, PyCaret, or similar ML frameworks.
• Experience in retraining or maintaining classification models in a production environment.
• Strong skills in experimental design and model evaluation.
• Experience in defining baselines, metrics, test sets, and error-analysis processes.
• Ability to assess model quality and demonstrate quantifiable improvements.
• In-depth understanding of scalability and ML inference costs.
• Strong English communication abilities.
• Nice-to-have: experience in entity resolution, record linkage, or deduplication.
• Nice-to-have: familiarity with ranking and similarity scoring methods.
• Nice-to-have: knowledge of retrieval, clustering, or candidate-generation techniques.
• Nice-to-have: LLM/embedding solutions designed for cost efficiency and scalability.
• Nice-to-have: experience with Spark, Snowflake, Databricks, or BigQuery.
• Nice-to-have: background in company, domain, website, or firmographic data.
• Nice-to-have: experience handling multilingual datasets.
• Strong analytical skills and an experimental mindset.
• Commitment to intellectual honesty and a willingness to report negative findings.
• High degree of autonomy and self-direction.
• Exceptional written and verbal communication skills.
• Ability to advocate for technical recommendations with stakeholders.
• Strong problem-solving capabilities.
• Comfort in navigating ambiguity and managing large-scale datasets.
• Ability to balance model quality, cost, and scalability effectively.
• Option for remote work.
• Full-time employment.
HighLevel
HighLevel
Brown and Caldwell
Get handpicked remote jobs straight to your inbox weekly.