
AI Research Peer Review Evaluator – ML/AI
Posted 4 days ago

Posted 4 days ago
This is a fully remote position, open to applicants in Switzerland.
• Analyze and interpret ML/AI research papers to grasp their fundamental contributions, methodologies, experimental setups, and assertions.
• Assess original human peer reviews to establish a reliable expert baseline for each publication.
• Compare AI-generated peer reviews against this baseline utilizing a structured scoring rubric.
• Evaluate the technical correctness, analytical depth, constructive value, and significance of each AI review.
• Detect hallucinations, unsupported assertions, overlooked technical issues, or valuable insights highlighted by AI reviewers.
• Conduct side-by-side comparisons of two AI-generated reviews to ascertain which one offers a more robust or beneficial analysis.
• Investigate and validate pertinent academic literature using platforms such as Google Scholar, arXiv, or Semantic Scholar.
• Confirm whether the referenced prior work was accessible prior to the paper’s submission date.
• Deliver concise, evidence-supported rationales for evaluation decisions and consistently apply the project rubric.
• Determine if agentic AI reviewers offer meaningful advantages over expert human reviewers.
• Possess a Master’s, PhD, or be actively pursuing graduate studies in Machine Learning, Artificial Intelligence, Computer Science, Statistics, or a closely aligned technical discipline.
• Have contributed to at least one scientific or research publication, preferably as the first author, though co-authors and significant contributors are also encouraged to apply.
• Experience in critically reviewing ML/AI research papers, including analysis of methodologies, experimental designs, outcomes, limitations, and scientific assertions.
• Familiarity with major ML/AI research venues, such as NeurIPS, ICML, ICLR, ACL, CVPR, or similar conferences and journals.
• Previous academic peer-review experience, ideally for an ML/AI conference or journal, is strongly preferred.
• Comfortable performing academic literature searches and verifying prior works, publication dates, citations, and claims of novelty.
• Possess strong analytical and written communication skills, with the ability to differentiate substantial technical concerns from superficial criticisms.
• Capable of providing clear, concise, evidence-based justifications for your evaluations.
• Able to consistently implement detailed evaluation guidelines and scoring rubrics across multiple papers and reviews.
• Exceptional attention to detail, especially when identifying factual inaccuracies or hallucinated technical claims.
• Fully remote and flexible — work from any location.
• Part-time contractor position with adaptable hours.
• Engage directly in the evaluation of innovative agentic AI systems for scientific research.
• Utilize your ML/AI research expertise to enhance the quality of AI-generated scientific peer reviews.
Gramian Consulting
Cotiviti
Netflix
SE Ranking
Get handpicked remote jobs straight to your inbox weekly.