
AI Safety Expert β English, Punjabi
Posted 14 hours ago

Posted 14 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments of conversational AI models and agents by utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data through the annotation of failures, classification of vulnerabilities, and identification of systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistency in testing processes.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs on sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Broaden evaluation coverage to reveal vulnerabilities that automated tests may overlook.
β’ Enhance the safety, robustness, and trustworthiness of customer AI systems.
β’ Native or fluent proficiency in English and Punjabi.
β’ Strong discernment regarding language and content; capable of evaluating the accuracy, completeness, and appropriateness of AI responses.
β’ Skill in identifying subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines, taxonomies, benchmarks, playbooks, and quality standards.
β’ Ability to articulate reasoning clearly to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects, task types, and client needs.
β’ Status as an independent contractor.
β’ H1-B and STEM OPT candidates are not eligible for consideration.
β’ Preferred experience in adversarial machine learning, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation probing, abuse analysis, conversational AI testing, psychology, acting, or unconventional adversarial writing.
β’ Fully remote position with flexible, self-managed scheduling.
β’ Weekly compensation via Stripe or Wise.
β’ Competitive salary.
β’ Access to wellness resources and clear guidelines for projects requiring higher sensitivity.
β’ Reasonable accommodations available upon request.
β’ Referral bonuses of up to $90 for each successful referral.
β’ Opportunity to collaborate with top researchers in the field.
β’ Gain experience in human data-driven AI red teaming at the cutting edge of safety.
Mercor
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.