
Lead Data Scientist
Posted Jun 17

Posted Jun 17
This is a fully remote position, open to applicants in New York.
• Gather, analyze, and interpret both small and large datasets to reveal valuable insights that aid in the development of statistical methods and machine learning algorithms.
• Oversee the design, training, and implementation of NLP and transformer-based models for applications in financial surveillance and supervision, such as detecting misconduct, market abuse, trade manipulation, and insider communication.
• Create machine learning models and conduct analytics in accordance with established workflows while actively seeking opportunities for optimization and improvement.
• Perform data annotation and conduct quality reviews.
• Engage in exploratory data analysis and analysis of model failure states.
• Contribute to model governance, documentation, and frameworks for explainability that adhere to both internal and regulatory AI standards.
• Provide guidance to clients and prospects during the fine-tuning and development processes of machine learning models and analytics.
• Mentor junior team members in model development and exploratory data analysis.
• Collaborate with Product Manager(s) to gather project and product requirements and translate these into technical tasks using the team’s tools, techniques, and procedures.
• Pursue ongoing self-directed personal development.
• In-depth knowledge of **financial markets, compliance, surveillance, supervision, or regulatory technology**.
• Experience with one or more data science and machine/deep learning frameworks and tools, such as scikit-learn, H2O, Keras, PyTorch, TensorFlow, pandas, NumPy, caret, and tidyverse.
• Strong grasp of data science and statistical principles (including regression, Bayes, time series, clustering, P/R, AUROC, exploratory data analysis, etc.).
• Comprehensive understanding of key programming concepts (e.g., split-apply-combine, data structures, object-oriented programming).
• Solid foundation in statistics (including hypothesis testing, ANOVA, chi-square tests, etc.).
• Familiarity with NLP transfer learning, including word embedding models (such as GloVe, fastText, word2vec) and transformer models (including BERT, SBERT, HuggingFace, and GPT-x, etc.).
• Experience using natural language processing toolkits such as NLTK, spaCy, and Nvidia NeMo.
• Knowledge of microservices architecture and continuous delivery concepts in machine learning, alongside related technologies like Helm, Docker, and Kubernetes.
• Understanding of Deep Learning techniques for NLP.
• Familiarity with LLMs, including using Ollama & Langchain.
• Excellent verbal and written communication skills.
• Proven ability to collaborate effectively and thrive in a team-oriented environment.
• **Preferred Qualifications**
• Master’s or Doctor of Philosophy degree in Computer Science, Applied Mathematics, Statistics, or a related scientific discipline.
• Familiarity with cloud computing platforms including AWS, GCS, and Azure.
• Experience with automated supervision, surveillance, and compliance tools.
• Comprehensive health and wellness programs.
• Opportunities for professional development and continuous learning.
• Flexible working arrangements to promote work-life balance.
• Collaborative and inclusive work environment.
• Competitive salary and performance-based incentives.
Leega
Knowtion Health
Moniepoint Inc. (Formerly TeamApt Inc.)
Hitachi
Get handpicked remote jobs straight to your inbox weekly.