
Member of Technical Staff, Frontier AI
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in New York.
• Lead research and evaluation initiatives from problem identification to data design, quality calibration, and signal validation.
• Establish rigorous methodologies to assess whether experimental results yield reliable and defensible research signals.
• Analyze model and system failures to uncover root causes, edge cases, and potential areas for enhancement.
• Assess datasets, experiments, and conclusions against relevant quality benchmarks.
• Serve as a quality checkpoint when signal strength, data integrity, or supporting evidence is lacking.
• Create ML-focused data systems, including task definitions, annotation schemas, rubrics, incentives, and supporting pipelines.
• Organize data and evaluation workflows around downstream model performance objectives.
• Convert ambiguous real-world behaviors into quantifiable evaluation frameworks and new data categories.
• Identify gaps in evaluation or dataset coverage and recommend further investment or iterations.
• Develop quality assurance processes that uphold consistent research standards.
• Examine model and system behaviors to detect recurring weaknesses and performance constraints.
• Iterate on evaluations, datasets, feedback loops, and quality criteria.
• Utilize experimental results to inform enhancements in model or agent performance.
• Decide when research directions should be expanded, revised, paused, or terminated based on evidence.
• Maintain a systems-level viewpoint focused on comprehensive AI performance.
• Collaborate with researchers, domain experts, operators, and cross-functional teams during project initiation, calibration, and iterations.
• Articulate research findings, trade-offs, limitations, and signal strength to both technical and non-technical stakeholders.
• Convert research advancements into evidence-based narratives.
• Facilitate alignment between experimental initiatives and real-world system demands.
• Provide technical judgment in ambiguous, high-impact research settings.
• Seasoned technical professional.
• Strong professional judgment regarding the quality of research signals.
• Experience in designing ML-oriented datasets, evaluation frameworks, annotation systems, rubrics, or QA processes.
• Capability to translate complex and ambiguous real-world system behaviors into structured research and evaluation opportunities.
• Strong sense of ownership and comfort in making decisions in uncertain or rapidly changing environments.
• Excellent written and verbal communication abilities.
• Skill in clearly explaining technical trade-offs, limitations, evidence quality, and research findings.
• Proven record of working directly with researchers, technical experts, or domain specialists during project calibration and iterations.
• Systems-level understanding of model, agent, or AI system performance.
• Experience with reinforcement learning environments, simulators, or feedback-driven training systems is advantageous.
• Beneficial experience in enhancing agentic or AI systems operating within real-world workflows.
• Prior experience in applied research or production environments with direct influence on deployed systems is advantageous.
• Strongly valued experience in designing evaluations for complex or real-world tasks.
• Familiarity with expert incentive design or high-stakes technical research programs is beneficial.
• Work must be completed without utilizing any confidential or proprietary information from any employer, client, institution, or third party.
• Fully remote work.
• Full-time engagement.
• Remote consulting opportunity.
• Compensation ranging from $600,000 to $2,000,000 per year.
Rich Products Australia
Coursedog
DaCodes.
DaCodes.
Get handpicked remote jobs straight to your inbox weekly.