
Data Science Engineer – ML, GenAI
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in Canada.
• Design, prototype, and assess LLM/NLP solutions, which encompass text classification, entity extraction, semantic search, summarization, and various NLP/AI applications.
• Convert successful prototypes into production-ready, clean, modular, maintainable, and thoroughly tested code.
• Implement object-oriented design and software engineering methodologies.
• Evaluate existing codebases, data pipelines, and applications to identify elements to preserve, enhance, consolidate, or redesign.
• Independently explore, prepare, transform, and analyze data using Databricks.
• Work collaboratively with data engineering and application development teams to integrate Databricks, data pipelines, models, and application services.
• Engage in technical discussions and contribute to architectural decisions.
• Establish development practices aimed at enhancing reliability, scalability, testing, and maintainability.
• Convert business requirements into actionable technical solutions in collaboration with engineering, data, and leadership teams.
• Advance AI and data science projects from initial experimentation to dependable production solutions.
• Develop scalable data pipelines, reusable machine learning capabilities, and maintainable software for Irth’s products and industries.
• 3–5 years of experience in data science, machine learning, or ML engineering, including production-quality code that goes beyond notebooks.
• Master’s degree or equivalent professional background in data science, machine learning, computer science, statistics, or a related field.
• Solid foundation in statistics and modeling methodologies.
• Practical experience with NLP and/or LLM-based solutions.
• Knowledge of Retrieval-Augmented Generation (RAG).
• Familiarity with embeddings and similarity search, including cosine similarity.
• Experience in hybrid search that combines BM25 and semantic search.
• Skills in text deduplication using fuzzy matching and MinHash.
• Proficiency in text classification and Named Entity Recognition (NER).
• Experience with search re-ranking.
• Expertise in prompt engineering.
• Strong command of Python, including libraries such as pandas, scikit-learn, PyTorch, and other relevant ML libraries.
• Knowledge of object-oriented programming, clean code practices, and software design principles.
• Proficient in statistical methods including exploratory data analysis, hypothesis testing, and statistical modeling.
• Hands-on experience with Databricks or a similar data/analytics platform.
• Understanding of the machine learning lifecycle, including training, evaluation, deployment, and monitoring for model drift and data quality.
• Experience in designing and developing REST APIs.
• Strong written and verbal communication skills.
• Nice-to-have: experience with large-scale unstructured text data, media monitoring, web scraping, or APIs.
• Nice-to-have: background in communications, public relations, media, or a related industry.
• Nice-to-have: familiarity with vector databases/stores, RAG frameworks, and NLP model fine-tuning.
• Nice-to-have: experience with MLflow or similar tools.
• Nice-to-have: knowledge of TypeScript and Node.js.
• Nice-to-have: experience with CI/CD and version control systems, including Git or GitHub Actions.
• Competitive compensation package commensurate with experience and qualifications.
• Medical, Dental, and Vision Insurance.
• 401(k) Plan with Company Match.
• Generous Paid Time Off (PTO).
• Company-Paid Holidays.
• Flexible Work Options – Opportunities for remote work are available, depending on the role and business requirements.
• On-Call Compensation – Additional pay for eligible on-call shifts.
InductiveHealth Informatics
Autodesk
Bixal
Mondelēz International
Get handpicked remote jobs straight to your inbox weekly.