
AI Safety Expert β English, Gujarati
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments of conversational AI models and agents by employing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation techniques.
β’ Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing methods.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs related to sensitive issues such as bias, misinformation, or harmful behaviors.
β’ Identify vulnerabilities that automated tests may overlook.
β’ Broaden evaluation coverage to minimize unexpected outcomes in production.
β’ Enhance customer AI systems through adversarial testing.
β’ Native fluency in both English and Gujarati is essential.
β’ Strong judgment regarding language and content is necessary.
β’ Capable of evaluating whether AI responses are accurate, complete, and appropriate, with the ability to articulate the reasoning behind assessments.
β’ Skilled at noticing subtle errors, inconsistencies, and gaps in AI outputs.
β’ Ability to consistently adhere to guidelines and maintain quality standards.
β’ Proficient in conveying reasoning clearly to both technical and non-technical audiences.
β’ Adaptable to various projects, task types, and client needs.
β’ Must operate as an independent contractor.
β’ Should not require H1-B or STEM OPT sponsorship.
β’ Nice-to-have: experience in adversarial machine learning, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Nice-to-have: background in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
β’ Nice-to-have: experience in socio-technical risk, including harassment/disinformation probing, abuse analysis, or testing conversational AI.
β’ Nice-to-have: expertise in psychology, acting, or writing that supports unconventional adversarial thinking.
β’ Fully remote position.
β’ Flexible working hours; tasks can be completed according to your own schedule.
β’ Weekly payments through Stripe or Wise based on services provided.
β’ Project durations may vary based on needs and performance, allowing for extensions, reductions, or early conclusions.
β’ Participation in higher-sensitivity projects is optional.
β’ Clear guidelines and wellness resources are available for work involving sensitive content.
β’ Reasonable accommodations are provided upon request.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Chance to contribute to making AI systems more robust, safe, and trustworthy.
β’ Competitive compensation.
β’ Collaboration with leading researchers in the field.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.