
MCP & AI Connector Evaluation Specialist
Posted Jul 27

Posted Jul 27
This is a fully remote position, open to applicants in New York.
• Develop realistic prompts and scenarios for intricate, high-context personal tasks.
• Create workflows across various domains including travel, health research, dining, activity planning, home services, career searches, and personal organization.
• Utilize MCP-enabled tools, plugins, and connectors to execute multi-step workflows.
• Test integrations involving platforms like Google Drive, Notion, travel services, and others.
• Assess whether AI systems appropriately select and utilize connected tools.
• Complete designated workflows while capturing the screen.
• Document actions taken, tools used, decisions made, successes, failures, overreach, missing context, and impractical results.
• Evaluate whether outputs are personalized, realistic, useful, safe, and logically sound.
• Provide written explanations regarding model strengths, weaknesses, and patterns of failure.
• Develop and consistently implement detailed scoring rubrics.
• Identify instances of incomplete reasoning, unrealistic recommendations, and improper tool usage.
• Advanced practical experience with MCP, plugins, and AI connectors.
• Regular use of connected LLM tools, ideally several times each week.
• Extensive personal use of AI for planning, research, organization, and decision-making.
• An active LLM account with approximately six months or more of consistent usage history.
• Experience leveraging AI for high-context, multi-step personal workflows.
• Strong writing skills, judgment, reasoning abilities, and attention to detail.
• Capability to articulate why AI outputs are effective, incomplete, unsafe, or unrealistic.
• Experience in designing and applying structured evaluation rubrics.
• Availability to contribute at least 20 hours per week.
• Ability to complete assigned tasks within approximately 24 hours.
• Current residence in the United States.
• While formal technical credentials are appreciated, extensive hands-on experience with LLM tools and connected workflows is prioritized.
• A desktop or laptop computer is required; Chromebooks are not supported.
• Screen recording is mandatory during task execution.
• Willingness to electronically sign a data-sharing consent form.
• More than 100 hours of prior experience in rubric design, evaluation, or quality assessment is a plus.
• Familiarity with Google Drive, Notion, travel platforms, and other connected applications is beneficial.
• A background in model evaluation, human data, user research, quality assurance, or AI training is advantageous.
• Weekly payments via Stripe or Wise.
• Fully remote work within the United States.
• Potential for ongoing work following successful completion of the initial trial period.
• Opportunity to gain experience in MCP, connector evaluation, rubric development, and personalized AI research.
• Projects may be extended, shortened, or modified based on scope and performance.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.