
AI Safety Expert β English, Gujarati
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team evaluations of conversational AI models and agents by employing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation techniques.
β’ Generate annotated human data by identifying failures, categorizing vulnerabilities, and highlighting systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
β’ Create reproducible reports, datasets, and attack case studies for clients.
β’ Analyze AI outputs related to sensitive topics, including bias, misinformation, and harmful behaviors.
β’ Contribute to the expansion of evaluation coverage and the enhancement of customer AI systems.
β’ Proficient fluency in both English and Gujarati is essential.
β’ Strong discernment regarding language and content.
β’ Capability to evaluate whether AI responses are accurate, complete, and suitable, with the ability to articulate the reasoning behind assessments.
β’ Attention to detail to identify subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines and quality standards.
β’ Skill in clearly articulating reasoning to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects, tasks, and clientele.
β’ Must operate as an independent contractor.
β’ Must not require H1-B or STEM OPT sponsorship/support.
β’ Preferred qualifications: experience in adversarial machine learning, including familiarity with jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Preferred qualifications: background in cybersecurity, including penetration testing, exploit development, or reverse engineering.
β’ Preferred qualifications: experience in socio-technical risk, encompassing harassment/disinformation probing, abuse analysis, or conversational AI testing.
β’ Preferred qualifications: experience in psychology, acting, or writing to foster unconventional adversarial thinking.
β’ Fully remote position.
β’ Flexible working hours; complete tasks according to your own schedule.
β’ Weekly compensation via Stripe or Wise.
β’ Competitive salary.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Collaboration with leading researchers in the field.
β’ Reasonable accommodations available upon request.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.