
AI Safety Experts β English, Assamese
Posted Sep 19

Posted Sep 19
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents through techniques such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by documenting failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Adhere to established taxonomies, benchmarks, and playbooks to ensure consistency in testing.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs related to sensitive subjects like bias, misinformation, or harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Broaden evaluation coverage and minimize unexpected issues in production.
β’ Enhance customer AI systems and contribute to the development of safer, more resilient, and trustworthy models.
β’ Native proficiency in both English and Assamese is essential.
β’ Strong judgment regarding language and content, including the ability to assess the accuracy, completeness, and suitability of AI responses.
β’ Capability to articulate reasoning clearly to both technical and non-technical audiences.
β’ Meticulous attention to detail regarding subtle errors, inconsistencies, and gaps.
β’ Ability to consistently adhere to guidelines and maintain quality standards.
β’ Flexibility to adapt across various projects, task types, and client needs.
β’ Preferred areas of expertise include adversarial ML, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation analysis, abuse investigation, conversational AI testing, psychology, acting, or writing.
β’ Must be engaged as an independent contractor.
β’ H1-B and STEM OPT candidates are not eligible for support.
β’ Fully remote work.
β’ Flexible working hours / ability to create your own schedule.
β’ Weekly payments through Stripe or Wise.
β’ Project timelines may be extended, shortened, or concluded early based on needs and performance.
β’ Access to wellness resources and clear guidelines for higher-sensitivity projects.
β’ Reasonable accommodations available upon request.
β’ Referral bonuses of up to $90 for each successful referral (restrictions may apply).
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.