
NLP Data Scientist
Posted Sep 23

Posted Sep 23
This is a fully remote position, open to applicants in India.
• Spearhead fine-tuning projects for LLMs utilizing PEFT methodologies such as LoRA, QLoRA, Prefix Tuning, and Prompt Tuning to enhance open-source frameworks like Llama and Mistral.
• Select, cleanse, and organize training datasets effectively.
• Develop automated processes for data deduplication, tokenization methods, synthetic data generation workflows, and human-in-the-loop validation systems.
• Establish reinforcement learning alignment layers through RLHF or DPO to ensure model safety, helpfulness, and maintain tone guidelines.
• Employ post-training quantization strategies including GGUF, AWQ, and GPTQ to reduce parameter loss and manage computing resources.
• Create evaluation benchmarks and metrics panels utilizing automated validation tests such as BLEU, ROUGE, and tailored verification matrices.
• Conduct audits on model hallucinations, factual accuracy, and domain alignment.
• Oversee distributed deep learning training tasks across multi-GPU computing clusters.
• Scale pipeline configurations, tensor parallelism settings, and gradient checkpointing scripts.
• Collaborate with MLOps infrastructure teams to prepare model weight checkpoints for scalable cloud deployment and real-time inference serving.
• Total experience: 4 to 8 years.
• A minimum of 3 years of focused natural language processing (NLP) and practical Large Language Model (LLM) fine-tuning experience.
• Required certification: Google Cloud Certified Professional Machine Learning Engineer or AWS Certified Machine Learning - Specialty.
• 4 to 8 years of fundamental data science or advanced machine learning engineering experience, with at least 3 dedicated years spent actively training, evaluating, and fine-tuning natural language processing systems.
• Expert-level proficiency in Python, deep learning frameworks (PyTorch), transformer architectures (Hugging Face Transformers, Accelerate, PEFT), and vector computations.
• In-depth structural knowledge of attention mechanisms, tokenization limitations, context window decay behaviors, loss function optimization, and hardware computing constraints (CUDA).
• Preferred: Master’s or Ph.D. in Computer Science, Data Science, Computational Linguistics, or a related quantitative discipline with research focused on neural network text models.
• Previous experience in implementing custom embedding model architectures or enhancing domain-specific classification layers within constrained enterprise environments.
• Competitive salary and performance-based bonuses.
• Comprehensive health, dental, and vision insurance.
• Flexible working hours and remote work options.
• Opportunities for professional development and continued learning.
• Collaborative and inclusive work environment.
Keyrus
CareDx, Inc.
Get handpicked remote jobs straight to your inbox weekly.