
AI Safety Expert, English β Vietnamese
Posted Sep 15

Posted Sep 15
This is a fully remote position, open to applicants in United States.
β’ Engage in red-teaming for conversational AI models and agents by executing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Generate human data through annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing processes.
β’ Create reproducible reports, datasets, and attack scenarios for our clients.
β’ Evaluate AI outputs that involve sensitive subjects like bias, misinformation, or harmful conduct.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Increase evaluation coverage while minimizing unexpected issues in production.
β’ Assist Mercor clients in enhancing the safety and reliability of their AI systems.
β’ Native fluency in both English and Vietnamese is a must.
β’ Prior experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical probing is essential.
β’ Capability to probe systems in an adversarial manner and test their limits.
β’ Familiarity with frameworks or benchmarks for systematic testing.
β’ Proficiency in articulating risks to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects and clients.
β’ Independent contractor status is required.
β’ Candidates must not hold H1-B or STEM OPT status.
β’ Nice-to-have: experience in adversarial machine learning, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Nice-to-have: background in cybersecurity encompassing penetration testing, exploit development, or reverse engineering.
β’ Nice-to-have: experience in socio-technical risk management, including harassment/disinformation probing, abuse analysis, or testing of conversational AI.
β’ Nice-to-have: experience in psychology, acting, or writing that fosters unconventional adversarial thinking.
β’ Fully remote position that allows for flexible scheduling.
β’ Receive weekly payments through Stripe or Wise based on services provided.
β’ Project timelines may be extended, shortened, or concluded early based on requirements and performance.
β’ Access to wellness resources and comprehensive guidelines for projects of higher sensitivity.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Competitive compensation package.
β’ Referral bonuses of up to $100 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.