
AI Safety Experts β English, Malay
Posted Aug 10

Posted Aug 10
This is a fully remote position, open to applicants in United States.
β’ Conduct red team assessments on conversational AI models and agents by utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation techniques.
β’ Create human data through the annotation of failures, classification of vulnerabilities, and identification of systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
β’ Generate reproducible reports, datasets, and attack scenarios for client use.
β’ Investigate AI systems to reveal vulnerabilities that automated tests may overlook.
β’ Broaden evaluation coverage and minimize unexpected issues during production.
β’ Assist Mercor clients in enhancing the safety and resilience of their AI systems.
β’ Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical probing.
β’ Capability to challenge systems to their limits through adversarial testing.
β’ Familiarity with frameworks or benchmarks instead of random hacking approaches.
β’ Proficiency in articulating risks clearly to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects and clientele.
β’ Native fluency in both English and Malay.
β’ Experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction is a plus.
β’ Background in cybersecurity specialties such as penetration testing, exploit development, or reverse engineering is a plus.
β’ Expertise in socio-technical risks including harassment/disinformation probing, abuse analysis, or testing of conversational AI is a plus.
β’ Creative probing experience in fields such as psychology, acting, or writing is a plus.
β’ Candidates must not be on an H1-B or STEM OPT visa.
β’ Fully remote position.
β’ Flexible scheduling options.
β’ Participation in higher-sensitivity projects is optional.
β’ Access to clear guidelines and wellness resources for projects involving sensitive content.
β’ Weekly payments made through Stripe or Wise.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Play a direct role in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Competitive compensation.
β’ Collaboration opportunities with leading researchers.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.