
Member of Technical Staff, Coding Research
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in New York.
β’ Design and manage evaluation frameworks for sophisticated coding agents.
β’ Create benchmark specifications, scoring methods, rubrics, and quality standards.
β’ Implement stringent methods for assessing coding-model performance across various software engineering tasks.
β’ Establish objective criteria for correctness, reasoning quality, robustness, and task completion.
β’ Uphold methodological rigor, reproducibility, and consistency throughout evaluation workflows.
β’ Generate high-quality datasets, exemplary cases, and structured evaluation protocols.
β’ Craft technical tasks for the reliable evaluation of cutting-edge coding systems.
β’ Develop data and evaluation workflows that support model development and iterative enhancements.
β’ Identify gaps in benchmark or dataset coverage and create new evaluation categories.
β’ Analyze coding-agent behavior to uncover systematic weaknesses, failure modes, and performance constraints.
β’ Investigate issues such as incorrect reasoning, implementation errors, tool-use failures, and incomplete task execution.
β’ Convert findings into actionable recommendations for model training and evaluation.
β’ Design experiments to test hypotheses regarding coding-model capabilities.
β’ Construct tools and infrastructure for large-scale experimentation, data generation, review processes, and evaluation pipelines.
β’ Automate technical procedures to enhance evaluation efficiency and research speed.
β’ Collaborate with researchers, engineers, and applied AI teams.
β’ Contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives.
β’ Effectively communicate complex technical findings to both specialized and general technical audiences.
β’ Strong software engineering background with proficiency in Python, C++, or similar programming languages.
β’ At least 3 years of experience in software engineering, machine learning, AI research, evaluation, or a related technical field.
β’ Experience in designing, reviewing, or validating technical assessments, benchmarks, coding tasks, or evaluation methodologies.
β’ Familiarity with large language models, coding agents, reinforcement learning, model evaluation, or related AI systems.
β’ Demonstrated ability to build tools, automate workflows, and enhance technical processes through systematic experimentation.
β’ Strong analytical abilities and capacity to investigate complex model behavior and technical failure modes.
β’ Excellent written and verbal communication skills.
β’ Capability to function effectively in fast-paced research environments characterized by significant ambiguity and shifting priorities.
β’ Experience with advanced AI systems, coding agents, or model-evaluation research is a plus.
β’ Expertise in designing benchmarks or datasets for large-scale machine learning systems is highly valued.
β’ Knowledge of agentic workflows, tool use, reinforcement learning, or post-training methodologies is beneficial.
β’ Publications, open-source contributions, or demonstrated technical leadership in AI, machine learning, or software engineering are advantageous.
β’ Work must be conducted without utilizing confidential or proprietary information from any employer, client, institution, or other third party.
β’ Fully remote work.
β’ Full-time engagement.
β’ Opportunity to contribute to pioneering AI research and development.
β’ Collaboration with researchers, engineers, and applied AI teams.
β’ Chance to contribute to technical reports, benchmark studies, research documentation, and external-facing research initiatives.
Rich Products Australia
Coursedog
DaCodes.
DaCodes.
Get handpicked remote jobs straight to your inbox weekly.