
Senior ML Engineer
Posted Jul 28

Posted Jul 28
This is a fully remote position, open to applicants in Colorado, +1 more state.
• Develop, implement, and sustain comprehensive ML training pipelines — covering everything from raw data ingestion and preprocessing to model training, evaluation, and deployment.
• Optimize large language models through techniques such as LoRA, QLoRA, and full fine-tuning; utilize PEFT strategies to find a balance between performance and computational costs.
• Execute and test reinforcement learning from human feedback (RLHF) workflows, incorporating PPO (Proximal Policy Optimization) and GRPO (Group Relative Policy Optimization) for aligning models and optimizing preferences.
• Host, deploy, and enhance LLMs in production environments using inference frameworks like vLLM, Text Generation Inference (TGI), Triton Inference Server, or ONNX Runtime.
• Assess, benchmark, and choose inference providers (e.g., Together AI, Fireworks, Groq, Replicate, AWS Bedrock, Azure OpenAI) based on trade-offs in latency, cost, throughput, and model capabilities.
• Create and maintain embedding pipelines — generate, index, and retrieve dense embeddings utilizing vector databases (Pinecone, pgvector, Weaviate, or similar) for RAG and semantic search functions.
• Implement and expose ML capabilities through the Model Context Protocol (MCP) — allowing AI agents to utilize model-backed tools in a structured and context-aware manner.
• Conduct thorough data analysis and processing: clean, transform, and curate datasets for training, fine-tuning, and evaluation; establish data quality and validation pipelines.
• Build comprehensive model evaluation frameworks — define metrics, create evaluation harnesses, conduct A/B testing, and monitor regressions across model versions.
• Collaborate with software engineers to integrate ML systems into product features via FastAPI services; ensure models are observable, versioned, and maintainable in production.
• Bachelor’s degree in computer science or statistics.
• 3–8+ years of practical ML engineering experience with a proven track record in production settings.
• Strong grasp of core ML principles: neural network architectures (transformers, attention mechanisms), loss functions, optimization algorithms, regularization, and model evaluation.
• Practical experience in fine-tuning LLMs (LoRA, QLoRA, PEFT, instruction tuning, DPO) on custom datasets using frameworks such as Hugging Face Transformers, TRL, or Axolotl.
• Direct experience with RL-based alignment methods — particularly PPO and GRPO — for reward modeling, preference optimization, and RLHF workflows.
• Experience in hosting and serving LLMs: vLLM, TGI, Triton, or similar; familiarity with model quantization (GPTQ, AWQ, int4/int8), batching strategies, and throughput enhancement.
• Working knowledge of major inference providers and cloud AI APIs; ability to evaluate and select vendors based on cost, latency, and capability criteria.
• Proficiency in embedding models (sentence-transformers, OpenAI embeddings, or equivalent) and vector search frameworks for RAG pipelines.
• Understanding of Model Context Protocol (MCP) and the ability to present ML functionalities as structured tools for agentic systems.
• Unlimited vacation for exempt employees.
• Paid holidays.
• Competitive medical, dental, and vision insurance for employees and their dependents.
• 401K retirement plan.
• Stock options.
• Company-paid life insurance.
• Health and flexible savings accounts.
• Reimbursements for cell phone, gym, and internet expenses.
• Paid parental leave.
• Tuition reimbursement.
• Employee Assistance Program (EAP).
• Free snacks (available in Denver and/or Fort Lauderdale).
• Engaging events (both virtual and in-person).
Doma
CSC Generation
Accelerant
Capgemini
Get handpicked remote jobs straight to your inbox weekly.