
Data Science Engineer β ML, GenAI
Posted Aug 12

Posted Aug 12
This is a fully remote position, open to applicants in Canada.
β’ Design, prototype, and assess LLM/NLP solutions, which encompass text classification, entity extraction, semantic search, summarization, and various NLP/AI applications.
β’ Convert successful prototypes into production-ready, clean, modular, maintainable, and thoroughly tested code.
β’ Implement object-oriented design and software engineering methodologies.
β’ Evaluate existing codebases, data pipelines, and applications to identify elements to preserve, enhance, consolidate, or redesign.
β’ Independently explore, prepare, transform, and analyze data using Databricks.
β’ Work collaboratively with data engineering and application development teams to integrate Databricks, data pipelines, models, and application services.
β’ Engage in technical discussions and contribute to architectural decisions.
β’ Establish development practices aimed at enhancing reliability, scalability, testing, and maintainability.
β’ Convert business requirements into actionable technical solutions in collaboration with engineering, data, and leadership teams.
β’ Advance AI and data science projects from initial experimentation to dependable production solutions.
β’ Develop scalable data pipelines, reusable machine learning capabilities, and maintainable software for Irthβs products and industries.
β’ 3β5 years of experience in data science, machine learning, or ML engineering, including production-quality code that goes beyond notebooks.
β’ Masterβs degree or equivalent professional background in data science, machine learning, computer science, statistics, or a related field.
β’ Solid foundation in statistics and modeling methodologies.
β’ Practical experience with NLP and/or LLM-based solutions.
β’ Knowledge of Retrieval-Augmented Generation (RAG).
β’ Familiarity with embeddings and similarity search, including cosine similarity.
β’ Experience in hybrid search that combines BM25 and semantic search.
β’ Skills in text deduplication using fuzzy matching and MinHash.
β’ Proficiency in text classification and Named Entity Recognition (NER).
β’ Experience with search re-ranking.
β’ Expertise in prompt engineering.
β’ Strong command of Python, including libraries such as pandas, scikit-learn, PyTorch, and other relevant ML libraries.
β’ Knowledge of object-oriented programming, clean code practices, and software design principles.
β’ Proficient in statistical methods including exploratory data analysis, hypothesis testing, and statistical modeling.
β’ Hands-on experience with Databricks or a similar data/analytics platform.
β’ Understanding of the machine learning lifecycle, including training, evaluation, deployment, and monitoring for model drift and data quality.
β’ Experience in designing and developing REST APIs.
β’ Strong written and verbal communication skills.
β’ Nice-to-have: experience with large-scale unstructured text data, media monitoring, web scraping, or APIs.
β’ Nice-to-have: background in communications, public relations, media, or a related industry.
β’ Nice-to-have: familiarity with vector databases/stores, RAG frameworks, and NLP model fine-tuning.
β’ Nice-to-have: experience with MLflow or similar tools.
β’ Nice-to-have: knowledge of TypeScript and Node.js.
β’ Nice-to-have: experience with CI/CD and version control systems, including Git or GitHub Actions.
β’ Competitive compensation package commensurate with experience and qualifications.
β’ Medical, Dental, and Vision Insurance.
β’ 401(k) Plan with Company Match.
β’ Generous Paid Time Off (PTO).
β’ Company-Paid Holidays.
β’ Flexible Work Options β Opportunities for remote work are available, depending on the role and business requirements.
β’ On-Call Compensation β Additional pay for eligible on-call shifts.
The Home Depot
HighLevel
HighLevel
TOPMIND
Get handpicked remote jobs straight to your inbox weekly.