
Member of Technical Staff – ML Platform
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in United States.
• Take ownership of the evaluation platform from start to finish, encompassing tools and systems to create, annotate, review, and modify evaluations at a frontier scale.
• Establish CLIs, APIs, GUIs, and storage solutions to ensure that evaluations are seamless, rapid, sophisticated, and collaborative.
• Collaborate directly with research teams focused on video, image, audio, agents, and robotics to comprehend measurement requirements and develop generalized platform solutions.
• Set organization-wide standards for model evaluation, covering reproducibility, metric definitions, reporting formats, and the reliability of results.
• Facilitate the adoption and integration of the platform across training, production model serving, and other research initiatives.
• Contribute to the ML Platform tools and systems that support Runway in training and serving frontier models.
• A minimum of 5 years of experience in constructing ML infrastructure or data platforms within production settings.
• Some experience with evaluation, experimentation, or benchmarking systems is required.
• Proficiency in Python and PyTorch is essential.
• Practical experience in managing large batch GPU workloads on Kubernetes.
• Experience in designing data pipelines and storage solutions for extensive media or model outputs.
• A keen focus on versioning and reproducibility.
• Familiarity with experimental statistics, including paired comparisons, confidence intervals, multiple-comparison issues, and inter-rater agreement.
• Comfortable building internal tools from end to end, spanning from the command line to the browser interface.
• Capability to lead a comprehensive technical area, gather requirements, set direction, make trade-offs, and drive a strategic roadmap.
• Knowledge of the complete model development lifecycle: data, training, evaluation, and serving.
• A self-starter who can work closely with research teams and act quickly.
• Strong systems thinking complemented by a pragmatic approach to production reliability.
• Humility and an open-minded attitude are essential.
• Nice to have: experience in developing evaluation suites for generative models.
• Nice to have: hands-on experience with LLM- or VLM-as-judge pipelines.
• Nice to have: familiarity with online experimentation platforms and linking offline metrics to product outcomes.
• Nice to have: prior experience evaluating agents or robotics policies.
• Competitive salary range of $240K–$290K for candidates located in the U.S.
• Access to top-tier models, agents, GPUs, storage, and cloud services.
• An equal opportunity workplace that is dedicated to diversity and inclusion.
Rich Products Australia
Coursedog
DaCodes.
DaCodes.
Get handpicked remote jobs straight to your inbox weekly.