
AI Safety Expert β English, Assamese
Posted Aug 30

Posted Aug 30
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human-generated data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing methodologies.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Analyze AI outputs related to sensitive issues such as bias, misinformation, and harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Broaden evaluation coverage to minimize unexpected challenges in production.
β’ Enhance customer AI systems and contribute to the development of safer, more reliable AI technologies.
β’ Native proficiency in English and Assamese.
β’ Previous experience in red teaming, particularly in AI adversarial contexts, cybersecurity, or socio-technical probing.
β’ Capability to test systems adversarially and push them to their limits.
β’ Familiarity with frameworks or benchmarks for structured testing approaches.
β’ Ability to communicate risks effectively to both technical and non-technical stakeholders.
β’ Flexibility to adapt to various projects and client needs.
β’ Desirable skills: adversarial ML, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation probing, abuse analysis, conversational AI testing, psychology, acting, or writing.
β’ Candidates must not be on an H1-B or STEM OPT visa.
β’ Fully remote position allowing for flexible scheduling.
β’ Weekly payments through Stripe or Wise based on services performed.
β’ Optional participation in high-sensitivity projects.
β’ Clear guidelines and wellness resources provided for sensitive-content projects.
β’ Competitive compensation.
β’ Referral bonuses of up to $90 for each successful referral, with no cap on the number of referrals (certain restrictions may apply).
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.