
AI Safety Experts, English, Kannada
Posted 16 hours ago

Posted 16 hours ago
This is a fully remote position, open to applicants in United States.
β’ Engage in red-team assessments of conversational AI models and agents by utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data through the annotation of failures, classification of vulnerabilities, and identification of systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistency in testing processes.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Analyze AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Broaden evaluation coverage and minimize unexpected issues in production.
β’ Enhance customer AI systems through adversarial testing.
β’ Native proficiency in both English and Kannada.
β’ Strong discernment regarding language and content, including assessing the accuracy, completeness, and appropriateness of AI responses.
β’ Capability to articulate reasoning clearly to both technical and non-technical audiences.
β’ Keen attention to subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines and quality standards.
β’ Flexibility to adapt across various projects, task types, and client needs.
β’ Independent contractor status is required.
β’ Candidates must not be on an H1-B or STEM OPT visa.
β’ Preferred qualifications include specialties in adversarial ML, cybersecurity, socio-technical risk, or creative probing.
β’ Experience in adversarial ML with jailbreak datasets, prompt injections, RLHF/DPO attacks, or model extraction is advantageous.
β’ Background in cybersecurity with penetration testing, exploit development, or reverse engineering is a plus.
β’ Experience in socio-technical risk involving harassment/disinformation probing, abuse analysis, or conversational AI testing is beneficial.
β’ Background in creative probing through psychology, acting, or writing is a plus.
β’ Fully remote position.
β’ Flexible working hours; tasks can be completed according to your own schedule.
β’ Weekly payments processed through Stripe or Wise.
β’ Project durations may vary based on needs and performance; they can be extended, shortened, or concluded early.
β’ Participation in higher-sensitivity projects is optional.
β’ Access to clear guidelines and wellness resources for dealing with sensitive content.
β’ Reasonable accommodations available upon request.
β’ Competitive compensation.
β’ Opportunity to collaborate with leading researchers in the field.
β’ Referral bonuses of up to $90 for each successful referral, subject to conditions.
Mercor
RWS Group
Mercor
Instacart
Get handpicked remote jobs straight to your inbox weekly.