
Member of Technical Staff – Post-Training
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in United States, +2 more countries.
• Create systems and methodologies for post-training of frontier models, which encompass supervised fine-tuning, reinforcement learning, preference optimization, reward modeling, and similar techniques.
• Convert open-ended research questions or partner requirements into hypotheses, experimental designs, evaluation strategies, and production-ready implementations.
• Develop and enhance evaluation frameworks, benchmarks, training environments, data-processing pipelines, and quality-control systems.
• Execute rapid iteration cycles: prototype, assess, analyze results, and apply insights to the next system or product.
• Collaborate with AI researchers and domain specialists to generate high-quality data, feedback, and evaluation methodologies.
• Recognize repeatable patterns across projects and transform them into reusable software and platforms.
• Elevate the technical standards through design judgment, effective communication, code quality, and mentorship.
• Contribute to benchmarks, open-source projects, research endeavors, and technical documentation where it provides leverage.
• Assist in shaping Handshake Labs' technical vision, operational culture, and reusable systems.
• A minimum of 3 years of proven experience in post-training, fine-tuning, or model-evaluation tasks.
• Relevant expertise may involve RL, SFT, LoRA/PEFT, full fine-tuning, RLHF, DPO, PPO, reward modeling, or training environments.
• Proficient in Python with the capability to write clean, efficient, and scalable code.
• Practical experience with contemporary ML tooling, especially PyTorch and large-scale data, training, or evaluation workflows.
• Strong experimental judgment, encompassing hypothesis formulation, metric selection, failure diagnosis, and differentiating signal from noise.
• Experience in designing systems and making trade-offs concerning quality, scalability, reliability, and reusability.
• Ability to thrive in a fast-paced, ambiguous environment with significant ownership.
• Strong collaborative communication skills and the ability to engage with researchers, engineers, domain experts, and customers.
• Particularly noteworthy: experience with large-scale ML training, inference, data, or evaluation systems.
• Especially compelling: work with LLM/agent benchmarks, evaluation methodologies, annotation systems, or data-quality frameworks.
• Especially noteworthy: involvement in reinforcement learning, alignment, model behavior, synthetic data, or human-in-the-loop systems.
• Especially compelling: published research, significant open-source contributions, or technical leadership in ML systems or AI research.
• Particularly attractive: translating research or repeated client projects into robust, reusable platforms.
• Equity in a rapidly-growing company.
• 401(k) matching.
• Competitive salary package.
• Financial coaching services.
• Paid parental leave.
• Fertility assistance.
• Parental coaching resources.
• Comprehensive medical, dental, and vision insurance.
• Mental health support services.
• $500 wellness stipend.
• $2,000 learning stipend.
• Ongoing professional development opportunities.
• Commuting assistance.
• Complimentary lunch.
• On-site gym facilities in the San Francisco office.
• Flexible paid time off.
• 15 holidays plus 2 flexible days.
• Team outings.
• Referral bonuses.
Silver.dev
Xentity Corporation
Get handpicked remote jobs straight to your inbox weekly.