
AI Safety Expert, English, Gujarati
Posted 13 hours ago

Posted 13 hours ago
This is a fully remote position, open to applicants in United States.
• Conduct red-team assessments of conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
• Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
• Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
• Generate reproducible reports, datasets, and attack scenarios.
• Evaluate AI outputs related to sensitive subjects like bias, misinformation, and harmful behaviors.
• Identify vulnerabilities that automated testing may overlook.
• Broaden evaluation scope and minimize unexpected issues in production.
• Enhance customer AI systems through adversarial testing strategies.
• Native proficiency in both English and Gujarati.
• Strong discernment regarding language and content quality.
• Capability to evaluate the accuracy, completeness, and appropriateness of AI responses, along with the ability to articulate reasoning.
• Meticulous attention to subtle errors, inconsistencies, and gaps in information.
• Consistent adherence to taxonomies, benchmarks, playbooks, guidelines, and quality standards.
• Ability to clearly convey reasoning to both technical and non-technical audiences.
• Flexibility to adapt across various projects, task types, and client requirements.
• Preferred expertise in adversarial machine learning, cybersecurity, socio-technical risk, or creative probing.
• Must be available as an independent contractor.
• H1-B and STEM OPT candidates will not be considered.
• Fully remote position.
• Flexible working hours to accommodate your schedule.
• Weekly compensation via Stripe or Wise.
• Project duration may vary based on needs and performance; extensions, reductions, or early conclusions are possible.
• Optional participation in higher-sensitivity projects.
• Clear guidelines and wellness resources available for handling sensitive content.
• Reasonable accommodations can be provided upon request.
• Opportunity to gain experience in human data-driven AI red teaming.
• Collaborate with leading researchers in the field.
• Referral bonuses of up to $90 for each successful referral, subject to certain limits.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.