
Python Engineer, AI Coding Agent Evaluator
Posted Sep 9

Posted Sep 9
This is a fully remote position, open to applicants in United States.
β’ Conduct comprehensive evaluations of AI-generated coding interactions from start to finish.
β’ Analyze the coherence of responses and the usefulness of preambles and reasoning.
β’ Assess whether outputs exhibit strong engineering judgment.
β’ Determine if interactions are suitable for seasoned developers.
β’ Evaluate the usefulness, overall correctness, and alignment with sound engineering principles.
β’ Assess explanations and reasoning in addition to the code itself.
β’ Differentiate between varying levels of response quality.
β’ Provide clear and decisive feedback on what was effective, what didn't work, and what seemed misleading.
β’ Assist in defining the characteristics of excellent interactions for tools like Cursor.
β’ Make informed yet subjective judgments regarding the behavior of AI coding agents.
β’ Highly skilled software engineer at a Senior+ level.
β’ Experience equivalent to a Staff / Principal-level engineer.
β’ Strong proficiency in TypeScript / JavaScript or Python.
β’ Practical experience with OpenAI Codex, Claude Code, and Cursor.
β’ Significant familiarity with contemporary AI-assisted development workflows.
β’ Capable of evaluating code without the necessity of executing or thoroughly reviewing every line.
β’ Comfortable providing direct, opinionated feedback.
β’ High standards for engineering quality.
β’ Complete a take-home evaluation exercise.
β’ Participate in one behavioral interview.
β’ Undergo a straightforward background check for the project.
β’ Required to choose preferred programming languages from the available list.
β’ Hourly rate between $100 and $200.
β’ Approximately 10 to 20 hours of work per week.
β’ Potential for extension beyond early May.
Gramian Consulting
Gramian Consulting
Intelance
Get handpicked remote jobs straight to your inbox weekly.