
AI Safety Expert β English, Tamil
Posted Sep 19

Posted Sep 19
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments of conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data by annotating failures, classifying vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs on sensitive topics, including bias, misinformation, or harmful behavior.
β’ Identify vulnerabilities that automated tests may overlook.
β’ Broaden evaluation coverage and mitigate surprises in production.
β’ Enhance customer AI systems and contribute to the development of safer, more reliable AI solutions.
β’ Proficiency in both English and Tamil at a native level is required.
β’ Strong discernment regarding language and content quality.
β’ Capability to evaluate whether AI responses are accurate, complete, and suitable, along with the ability to articulate reasons for assessments.
β’ Keen eye for subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines and quality standards is essential.
β’ Ability to communicate reasoning clearly to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects, task types, and clients.
β’ Must hold independent contractor status.
β’ Unable to accommodate H1-B or STEM OPT candidates.
β’ Preferred: experience in adversarial machine learning, including familiarity with jailbreak datasets, prompt injections, RLHF/DPO attacks, or model extraction.
β’ Preferred: background in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
β’ Preferred: experience in socio-technical risk management, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
β’ Preferred: creative probing experience in psychology, acting, or writing.
β’ Fully remote position.
β’ Flexible scheduling options.
β’ Weekly payments processed through Stripe or Wise.
β’ Optional involvement in higher-sensitivity projects.
β’ Clear guidelines and wellness resources available for projects involving sensitive content.
β’ Competitive compensation.
β’ Referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.