
AI Safety Expert β English, Kannada
Posted 12 hours ago

Posted 12 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents, utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation techniques.
β’ Create human data by annotating failures, categorizing vulnerabilities, and highlighting systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs related to sensitive subjects such as bias, misinformation, or harmful conduct.
β’ Identify vulnerabilities that automated tests may overlook.
β’ Broaden evaluation coverage and minimize unforeseen issues in production.
β’ Assist Mercor clients in enhancing the safety, robustness, and reliability of their AI systems.
β’ Native proficiency in both English and Kannada.
β’ Strong judgment regarding language and content, including the ability to assess whether AI responses are accurate, complete, and suitable.
β’ Capability to detect subtle errors, inconsistencies, and gaps.
β’ Consistently adhere to guidelines and quality standards.
β’ Ability to clearly articulate reasoning to both technical and non-technical audiences.
β’ Flexibility to adapt to various projects, task types, and client needs.
β’ Must operate as an independent contractor.
β’ Must not hold H1-B or STEM OPT status.
β’ Preferred: experience in adversarial ML, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Preferred: background in cybersecurity, encompassing penetration testing, exploit development, or reverse engineering.
β’ Preferred: experience with socio-technical risks, including harassment/disinformation probing, abuse analysis, or testing conversational AI.
β’ Preferred: creative probing skills in psychology, acting, or writing.
β’ Fully remote position that allows for a flexible schedule.
β’ Weekly payments through Stripe or Wise based on services rendered.
β’ Participation in higher-sensitivity projects is optional.
β’ Access to clear guidelines and wellness resources for sensitive-content projects.
β’ Competitive compensation.
β’ Direct involvement in human data-driven AI red teaming.
β’ Opportunity to collaborate with top researchers in the field.
β’ Referral bonuses of up to $90 for each successful referral, subject to limitations.
Gartner
Mercor
manara
Escalent
Get handpicked remote jobs straight to your inbox weekly.