
Senior Evaluation Algorithm Engineer
Posted 3 days ago

Posted 3 days ago
This is a fully remote position, open to applicants in Hong Kong, +2 more countries.
β’ Develop comprehensive LLM evaluation strategies for dialogue, financial trading, and various business contexts.
β’ Create evaluation metric frameworks and rubrics that yield measurable, reproducible, and comprehensible conclusions.
β’ Oversee the design and creation of evaluation datasets.
β’ Specify evaluation criteria and ensure scenario coverage.
β’ Set data annotation standards and maintain quality control procedures.
β’ Establish benchmarks that align with business requirements and possess discriminative capabilities.
β’ Examine model performance limits and identify failure modes.
β’ Offer actionable recommendations for improvement and partner with algorithm and product teams to facilitate model enhancements.
β’ Streamline and automate evaluation workflows.
β’ Develop sustainable evaluation platforms and toolchains.
β’ Collaborate with algorithm, product, and data teams to convert business and model goals into evaluation criteria.
β’ Transform evaluation insights into actionable R&D pathways and oversee their execution.
β’ A Master's degree or higher in Computer Science, Artificial Intelligence, Mathematics, Statistics, or relevant disciplines.
β’ A robust understanding of algorithmic principles and LLM methodologies, including training and fine-tuning processes.
β’ Practical LLM evaluation experience in a large technology organization.
β’ Engagement in commercial deployment evaluations rather than solely academic or offline benchmarks.
β’ Knowledge of the entire pipeline from evaluation data preparation and rubric formulation to evaluation-driven R&D.
β’ Experience with human evaluations, model-based automatic evaluations / LLM-as-a-judge, and metric calculations.
β’ Capability to delineate suitable evaluation criteria for various business scenarios.
β’ Proficiency in crafting clear, actionable, and effective rubrics.
β’ Systematic management of evaluation data representation, annotation consistency, and result dependability.
β’ Strong programming skills in Python.
β’ Background in automating evaluation workflows, constructing benchmarks, or developing evaluation platforms.
β’ Ability to manage data processing, evaluation script creation, and result analysis independently.
β’ Strong business acumen and effective communication skills.
β’ Aptitude for translating evaluation results into clear improvement strategies and fostering cross-team collaboration.
β’ Experience in evaluating dialogue systems, AI agents, or financial/trading LLMs is a plus.
β’ Experience in creating high-quality AI training/evaluation data or data annotation systems is advantageous.
β’ Familiarity with RLHF, reward models, or preference data-related tasks is a bonus.
β’ Competitive salary and comprehensive company benefits.
β’ Flexible work-from-home arrangements (may vary based on the business team's needs).
β’ Opportunities for career advancement and ongoing learning.
β’ Collaborate with top-tier talent in a user-focused global organization with a flat hierarchy.
β’ Engage in unique, fast-paced projects with a high degree of autonomy in an innovative setting.
β’ Results-oriented work environment.
Spyrosoft
Horizon3.ai
Jamf
Qualus
Get handpicked remote jobs straight to your inbox weekly.