
AI Safety Experts β English, Punjabi
Posted 13 hours ago

Posted 13 hours ago
This is a fully remote position, open to applicants in United States.
β’ Engage in red-teaming of conversational AI models and agents through jailbreak techniques, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by documenting failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Adhere to established taxonomies, benchmarks, and playbooks to ensure consistent testing.
β’ Develop reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Broaden evaluation coverage and minimize surprises in production.
β’ Enhance customer AI systems through adversarial testing.
β’ Must possess native fluency in both English and Punjabi.
β’ Strong discernment regarding language and content is essential.
β’ Capability to evaluate the accuracy, completeness, and appropriateness of AI responses, with the ability to articulate reasoning.
β’ Keen attention to detail for spotting subtle errors, inconsistencies, and omissions.
β’ Consistent adherence to guidelines and quality standards is necessary.
β’ Proficient in clearly conveying reasoning to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects, task types, and client needs.
β’ Must be an independent contractor.
β’ Candidates must not be on an H1-B or STEM OPT visa.
β’ Preferred: experience in adversarial machine learning, including familiarity with jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Preferred: background in cybersecurity, such as penetration testing, exploit development, or reverse engineering.
β’ Preferred: experience in socio-technical risks, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
β’ Preferred: skills in psychology, acting, or writing for unconventional adversarial perspectives.
β’ Fully remote position.
β’ Flexible working hours allowing you to set your own schedule.
β’ Weekly payment options via Stripe or Wise.
β’ Projects may be adjusted in length or concluded early based on needs and performance.
β’ Participation in higher-sensitivity projects is optional.
β’ Access to clear content guidelines and wellness resources.
β’ Competitive compensation.
β’ Reasonable accommodations available upon request.
β’ Referral bonuses of up to $90 for each successful referral (certain limits apply).
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.