
Research Data Scientist
Posted 8 hours ago

Posted 8 hours ago
This is a fully remote position, open to applicants in United States.
• Conduct both independent and collaborative research in the fields of Generative AI, LLMs, NLP, multimodal AI, machine learning, model evaluation, and AI data.
• Formulate research inquiries and convert complex AI/ML challenges into structured research methodologies and experimental designs.
• Design, implement, and analyze experiments aimed at assessing and enhancing AI/ML models and solutions.
• Create analytical models, prototypes, and research pipelines utilizing Python along with relevant ML frameworks.
• Develop and execute LLM evaluation frameworks, benchmarks, datasets, and evaluation criteria.
• Assess models for accuracy, robustness, bias, hallucination, reasoning, relevance, response quality, and various other performance metrics.
• Conduct model benchmarking, error analysis, comparative studies, and performance evaluations.
• Engage in RAG, SFT, RLHF/DPO, prompt engineering, fine-tuning, embeddings, and LLM optimization where applicable.
• Gather, clean, analyze, and interpret extensive and intricate structured and unstructured datasets.
• Execute EDA, statistical analysis, hypothesis testing, significance testing, correlation analysis, sampling, and error analysis.
• Develop and assess datasets, sampling methodologies, taxonomies, annotation frameworks, data quality frameworks, and evaluation criteria.
• Analyze data quality and pinpoint issues impacting model performance.
• Collaborate with annotation, data engineering, AI/ML, research, domain experts, and delivery teams.
• Contribute to research papers, technical reports, whitepapers, patents, benchmarks, internal publications, and additional research outputs.
• Present research outcomes, analytical insights, and technical recommendations to senior technical stakeholders.
• Engage in client-facing technical discussions and presentations as necessary.
• Transform business requirements into AI/ML solutions and convert complex research concepts into actionable recommendations.
• Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Statistics, Mathematics, Computational Science, or a related field.
• 4–7 years of practical research experience in AI/ML, Data Science, NLP, Generative AI, LLMs, or associated disciplines.
• Proven research experience with the ability to independently develop research questions, design experiments, analyze results, and communicate findings effectively.
• Demonstrated research credentials through publications, patents, conference presentations, open-source contributions, or significant AI/ML research initiatives.
• Strong proficiency in Python and SQL.
• Extensive hands-on experience with NumPy, Pandas, Scikit-learn, and preferably PyTorch/TensorFlow.
• In-depth understanding of machine learning algorithms, statistical methods, experimentation, data analysis, feature engineering, model evaluation metrics, hypothesis testing, and statistical inference.
• Practical exposure to LLMs, NLP, Generative AI, and multimodal AI.
• Experience with one or more of RAG, LLM evaluation, prompt engineering, fine-tuning, SFT, RLHF/DPO, embeddings, or model benchmarking.
• Experience handling large-scale structured and unstructured datasets.
• Familiarity with Git and cloud platforms such as AWS, Azure, or GCP is desirable.
• Bachelor’s/Master’s degree from IITs, NITs, or other top-tier engineering/research institutions is highly preferred.
• Preference will be given to candidates with publications in reputable conferences/journals and a robust academic/research profile.
• Comprehensive health and wellness programs.
• Opportunities for professional development and continuous learning.
• Collaborative work environment with access to cutting-edge technology.
• Flexible work arrangements and work-life balance.
Sigma Software Group
Get handpicked remote jobs straight to your inbox weekly.