
AI Developer
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
• Train and fine-tune Large Language Models (LLMs) utilizing supervised fine-tuning (SFT)
• Collaborate with open-source models like LLaMA, Mistral, Qwen, and similar architectures
• Create LoRA/Q-LoRA pipelines for efficient fine-tuning processes
• Develop and enhance data preprocessing workflows, encompassing tokenization and long-context management
• Utilize and extend Hugging Face Transformers & Datasets for both training and inference tasks
• Analyze and process structured and semi-structured data formats, including XML/XSD files
• Implement document parsing solutions for Office formats using python-docx and OpenXML
• Design and execute end-to-end Retrieval-Augmented Generation (RAG) pipelines for document-grounded question answering and knowledge retrieval
• Construct and sustain vector stores and embedding pipelines with technologies such as FAISS, Chroma, Weaviate, or pgvector
• Optimize retrieval strategies, including hybrid search, re-ranking, and chunking for domain-specific datasets
• Develop and manage Model Control Plane (MCP) server integrations to enable LLM access to tools, APIs, and external data sources
• Design agentic workflows utilizing MCP to ensure controlled, auditable access to internal systems and context
• Deploy, operate, and maintain models in fully offline and air-gapped environments
• Conduct model optimization and quantization using GGUF, GPTQ, AWQ, and bitsandbytes
• Build and sustain inference systems employing vLLM, TGI, and Ollama
• Enhance GPU utilization through CUDA, cuDNN, and VRAM-aware batching techniques
• Manage local CI/CD pipelines for machine learning models without reliance on cloud services
• Oversee local model registries, versioning, and artifacts management
• Ensure RAG and MCP components function effectively in offline and restricted network environments
• Develop Python backend services for machine learning training and inference workflows
• Collaborate with relational and vector databases for RAG storage solutions
• Utilize Docker and Git for development and deployment workflows
• Employ Azure DevOps for CI/CD processes, including local runners when necessary
• Strong proficiency in Python for backend and machine learning development
• Expertise in PyTorch or TensorFlow, along with scikit-learn and pandas
• Solid understanding of Postgres or MySQL databases
• Experience with Docker and Git version control
• Practical experience in LLM training, fine-tuning, and optimization techniques
• Familiarity with Hugging Face Transformers & Datasets
• Knowledge of XML/XSD and tools for parsing Office documents
• Experience in deploying models using vLLM, TGI, or Ollama
• Understanding of model quantization methods like GGUF, GPTQ, or AWQ
• Experience in GPU optimization and leveraging the CUDA ecosystem
• Proven experience in creating solutions for offline, on-premises, and air-gapped environments
• Practical experience in designing and implementing RAG pipelines, including embedding models, vector stores, and retrieval optimization
• Experience in building or integrating MCP servers
• Familiarity with advanced RAG techniques such as HyDE or multi-hop retrieval
• Ability to discuss complex technical topics with both technical and non-technical stakeholders
• Competitive salary and performance-based bonuses
• Flexible working hours and remote work options
• Continuous learning and professional development opportunities
• Access to cutting-edge tools and technologies
• Collaborative and inclusive work culture
Blend360
Get handpicked remote jobs straight to your inbox weekly.