
Philosophy Research & AI Benchmark Specialist
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in New York.
• Create original and thought-provoking multiple-choice philosophy questions that assess deep conceptual understanding.
• Craft self-contained, clear, and precise questions across various difficulty levels, including undergraduate, advanced undergraduate, and postgraduate.
• Assess pre-existing questions for accuracy, clarity, rigor, completeness, ambiguity, conceptual errors, and solvability.
• Revise questions and answer options, ensuring that chosen answers are justifiable.
• Record significant modifications and the rationale behind them.
• Formulate structured solution rationales utilizing rigorous philosophical argumentation and formal reasoning.
• Clarify distinctions between correct answers and plausible alternatives.
• Generate content related to AI ethics, responsible technology, epistemology, philosophy of technology, robotics, and artificial intelligence.
• Create and assess questions concerning formal ontology, knowledge representation, philosophy of science, explanation, methodology, evidence, and scientific reasoning.
• Provide 1–5 academic references per question when necessary.
• Evaluate question difficulty and assist in maintaining consistent academic quality across benchmark content.
• Contribute to benchmark materials used for evaluating advanced AI capabilities.
• Complete assigned tasks related to question authorship or verification independently.
• PhD or current doctoral candidacy in Philosophy or a closely related field.
• A Master's degree may be considered for candidates with exceptional expertise in a relevant philosophical subdomain.
• Strong proficiency in philosophical argumentation and analytical reasoning.
• Expertise in one or more areas such as Formal Ontology, Knowledge Representation, AI Ethics, Applied Epistemology, Philosophy of Technology, Robotics, or Philosophy of Science.
• Comprehensive knowledge of formal logic and pertinent philosophical literature.
• Ability to create rigorous, discerning multiple-choice assessments.
• Excellent written English skills and the capacity to convey complex ideas clearly and accurately.
• Academic research publications or university-level teaching experience are highly regarded.
• Experience in assessment design, examination development, or academic content review is advantageous.
• Strong attention to conceptual precision, ambiguity, and argumentative soundness.
• Expected availability of at least 10 hours per week.
• Work must be performed without utilizing confidential or proprietary information from any employer, client, institution, or third party.
• H1-B and STEM OPT support is not available for this position.
• Flexible scheduling based on project needs.
• Fully remote and asynchronous work environment.
• Part-time independent contractor arrangement.
• Opportunity to contribute to high-quality evaluation materials used to assess advanced AI capabilities.
CrowdStrike
Tempus AI
ŌURA
Vantor
Get handpicked remote jobs straight to your inbox weekly.