
AI Safety Expert β English, Urdu
Posted Sep 19

Posted Sep 19
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents through techniques such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by annotating failures, classifying vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing.
β’ Generate reproducible reports, datasets, and attack scenarios.
β’ Examine AI outputs for bias, misinformation, or harmful behaviors.
β’ Broaden evaluation coverage and pinpoint vulnerabilities that automated tests may overlook.
β’ Enhance customer AI systems by delivering actionable red-team artifacts.
β’ Native fluency in both English and Urdu is essential.
β’ Strong judgment regarding language and content is necessary.
β’ Capability to evaluate AI responses for accuracy, completeness, and appropriateness, along with the ability to articulate reasoning.
β’ Keen attention to detail for identifying subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines and quality standards.
β’ Ability to communicate reasoning clearly to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects, task types, and client needs.
β’ Must hold independent contractor status.
β’ H1-B or STEM OPT candidates will not be considered.
β’ Fully remote position.
β’ Flexible working hours, allowing you to complete tasks on your own schedule.
β’ Weekly payments via Stripe or Wise.
β’ Access to wellness resources and clear guidelines for projects requiring higher sensitivity.
β’ Competitive compensation.
β’ Opportunity to collaborate with leading researchers in the field.
β’ Referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.