
AI Interaction Evaluator β Codex, Claude Code
Posted 15 hours ago

Posted 15 hours ago
This is a fully remote position, open to applicants in United States, +13 more states.
β’ Assess the complete cycle of AI-generated coding interactions.
β’ Evaluate the usefulness and accuracy of outputs at a high level, ensuring they align with the thought process of a competent engineer.
β’ Analyze the quality of explanations and reasoning provided, not limited to just the code itself.
β’ Differentiate between various levels of response quality.
β’ Deliver clear and opinionated feedback regarding what was effective, what fell short, and what seemed inaccurate or misleading.
β’ Contribute to defining the standards of excellence in interactions with tools like Cursor.
β’ Determine if the responses, preambles, reasoning, and outputs from AI coding agents reflect strong engineering judgment.
β’ Engage in a take-home evaluation task and a behavioral interview.
β’ Highly experienced software engineer at the Senior+ level.
β’ Staff- or Principal-level engineer, or possess equivalent experience.
β’ Strong expertise in TypeScript/JavaScript or Python.
β’ Practical experience with OpenAI Codex, Claude Code, and Cursor.
β’ Profound understanding of contemporary AI-assisted development workflows.
β’ Capability to assess code without the need to execute or meticulously review every line.
β’ Comfortable providing direct and opinionated feedback.
β’ Maintain high standards for engineering quality.
β’ Able to make subjective yet rigorous judgments.
β’ Potential for extension beyond early May.
β’ Flexible part-time commitment of roughly 10β20 hours per week.
Progressive Leasing
apna
apna
Texas Research International
Get handpicked remote jobs straight to your inbox weekly.