
AI Safety Experts, English, Punjabi
Posted 4 hours ago

Posted 4 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team evaluations of conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Analyze AI-generated outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Document failures, categorize vulnerabilities, and identify systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to maintain consistency in testing.
β’ Generate reproducible reports, datasets, and attack case studies for clients.
β’ Investigate AI systems to reveal vulnerabilities that automated tests may overlook.
β’ Enhance evaluation coverage to minimize unexpected issues in production.
β’ Fortify customer AI systems through data-driven human red teaming.
β’ Fluent/native proficiency in English and Punjabi is essential.
β’ Strong judgment regarding language and content quality.
β’ Capable of evaluating AI responses for accuracy, completeness, and appropriateness, with the ability to articulate reasoning.
β’ Meticulous attention to detail regarding subtle errors, inconsistencies, and gaps.
β’ Consistently follow guidelines and uphold quality standards.
β’ Ability to clearly explain reasoning to both technical and non-technical audiences.
β’ Flexible and adaptable across various projects, task types, and customer needs.
β’ Engagement as an independent contractor.
β’ Must not require H1-B or STEM OPT support.
β’ Preferred specialties include adversarial ML, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation probing, abuse analysis, conversational AI evaluation, psychology, acting, and writing.
β’ Fully remote position.
β’ Flexibility to work on your own schedule.
β’ Weekly payments through Stripe or Wise based on services rendered.
β’ Project durations may be extended, shortened, or concluded early based on requirements and performance.
β’ Access to wellness resources and clear guidelines for higher-sensitivity projects.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Play a direct role in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Competitive compensation.
β’ Collaborate with leading researchers in the field.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.