
AI Interaction Evaluator, Codex / Claude Code
Posted Jul 25

Posted Jul 25
This is a fully remote position, open to applicants in Florida.
• Conduct comprehensive evaluations of AI-generated coding interactions from start to finish.
• Assess the outputs based on the following criteria:
• - Usefulness
• - Accuracy (at a high level)
• - Alignment with the thought processes of a proficient engineer.
• Evaluate the **quality of explanations and reasoning**, not solely the code.
• Differentiate between various levels of response quality (e.g., identifying what distinguishes a *2 from a 4*).
• Deliver clear and constructive feedback on:
• - What was effective
• - What was ineffective
• - What seemed “off” or misleading.
• Assist in defining what *excellence* looks like when engaging with tools like Cursor.
• Staff / Principal-level engineer (or equivalent experience).
• Strong expertise in one or more of the following:
• - TypeScript / JavaScript
• - Python
• Practical experience with:
• - OpenAI Codex
• - Claude Code
• - Cursor
• In-depth knowledge of modern AI-enhanced development workflows.
• Ability to assess code **without the necessity of executing or thoroughly reviewing every line**.
• Comfortable providing **direct and opinionated feedback**.
• High standards for what constitutes “good engineering.”
• Experience with tools like Cursor or similar AI-centric IDEs (Nice to Have).
• Previous exposure to prompt design or evaluation processes (Nice to Have).
• Experience in mentoring senior engineers or establishing engineering standards (Nice to Have).
• Competitive salary and performance-based bonuses.
• Flexible working hours and remote work options.
• Opportunities for professional development and growth.
• Collaborative and inclusive work environment.
Julesetmoi
National University
MeridianLink
Get handpicked remote jobs straight to your inbox weekly.