
AI Safety Experts – English, Thai
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
• Conduct red team assessments on conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Generate human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistency in testing.
• Create reproducible reports, datasets, and attack scenarios that customers can utilize.
• Discover vulnerabilities that automated tests often overlook.
• Broaden evaluation coverage to minimize unforeseen issues in production.
• Enhance customer AI systems and foster trust in their safety.
• Collaborate across various projects and clients for Mercor, which partners with AI labs and businesses to develop cutting-edge models.
• Previous experience in red teaming focused on AI adversarial tasks, cybersecurity, or socio-technical probing.
• Fluent in both English and Thai.
• Familiarity with jailbreaks, prompt injections, misuse cases, bias exploitation, or multi-turn manipulations.
• Capability to probe systems adversarially and challenge them to their limits.
• Proficiency in utilizing frameworks or benchmarks instead of employing random hacks.
• Ability to clearly articulate risks to both technical and non-technical audiences.
• Adaptability across various projects and clients.
• Preferred areas of expertise include adversarial machine learning, cybersecurity, socio-technical risks, or innovative probing techniques.
• Knowledge in adversarial ML, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction is advantageous.
• Understanding of cybersecurity concepts such as penetration testing, exploit development, or reverse engineering is a plus.
• Experience in socio-technical risk areas like harassment/disinformation probing, abuse analysis, or testing of conversational AI is beneficial.
• Background in creative probing through psychology, acting, or writing is an additional asset.
• Must be eligible to work without H1-B or STEM OPT support.
• Fully remote position.
• Flexible working hours allowing you to set your own schedule.
• Weekly payments through Stripe or Wise.
• Participation in higher-sensitivity projects is optional.
• Clear guidelines and wellness resources available for projects involving sensitive content.
• Opportunity to gain experience in human data-driven AI red teaming.
• Contribute to enhancing the robustness, safety, and trustworthiness of AI systems.
• Collaborate with top-tier researchers.
• Competitive compensation.
• Reasonable accommodations available upon request.
Progressive Leasing
apna
apna
Texas Research International
Get handpicked remote jobs straight to your inbox weekly.