
Data Scientist – NLP
Posted Jun 24

Posted Jun 24
This is a fully remote position, open to applicants in United States.
• Data Pre-processing - Showcase the ability to gather, cleanse, and prepare data sets for integration into computational models using Python.
• Ideal candidates will articulate various techniques they have utilized, including common pre-processing functions like stop word removal, stemming, lemmatization, and tokenization.
• Feature Engineering and Attribute Evaluation - Applicants must exhibit experience with NLP feature engineering techniques such as TF-IDF, word2vec, GloVe, and FastText, identifying crucial elements for modeling within the business process and existing data sets, as well as selecting evaluation protocols (model techniques).
• Modeling - Candidates should have hands-on experience in selecting classification modeling techniques appropriate for the business problem, including methodologies like machine learning (ML), supervised and unsupervised learning, regression, neural networks and deep learning, and natural language processing.
• Validation - Strong candidates will detail their experience in investigating, reporting, and justifying model outcomes.
• Visualization - Demonstrated experience in presenting the results of modeling activities, illustrating the insights gained, and explaining the significance of their findings in relation to the organization’s business challenges.
• A Master's degree is required, while a PhD is preferred in Statistics, Mathematics, Computer Science, or a related field.
• Extensive experience in utilizing SAS, R, or Python to support NLP applications such as Document Summarization, Named Entity Recognition, Sentiment Analysis, and/or Topic Modeling.
• A minimum of four years of experience in developing scalable, production-ready NLP solutions using sci-kit learn, Keras, TensorFlow, PyTorch, or Spark NLP.
• Proficient in using git/github for version control of source code.
• Experience in leveraging transformer architecture to create NLP models.
• Familiarity with open-source NLP libraries such as Gensim, SpaCy, or NLTK.
• Experience with models like BERT, GPT-J, RoBERTa, T5, or other transformer architectures.
• Experience with GenAI and Prompt Engineering is advantageous.
• Familiarity with Databricks and MLFlow is a plus.
• Experience in machine translation and transcription of foreign language documents using Microsoft Azure translation services is beneficial.
• Experience working within an AWS cloud environment, utilizing related AWS services such as Bedrock and Textract.
• Ability to coordinate and maintain user stories.
• Must be a US citizen.
• Must be able to obtain and maintain a Public Trust security clearance.
• Competitive salary with potential for bonuses.
• Employer-sponsored health care coverage.
• Funding for training and professional development.
• 401k matching contributions.
Solar Coca-Cola
Valtech
Nebius Group
Get handpicked remote jobs straight to your inbox weekly.