
AI Safety Experts, English, Assamese
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments of conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by annotating errors, categorizing vulnerabilities, and identifying systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing procedures.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs regarding sensitive issues including bias, misinformation, and harmful behaviors.
β’ Detect vulnerabilities that automated tests may overlook.
β’ Broaden evaluation coverage and minimize unexpected production issues.
β’ Collaborate on initiatives aimed at training and improving cutting-edge AI systems.
β’ Native proficiency in both English and Assamese is essential.
β’ Strong discernment regarding language and content, particularly in assessing the accuracy, completeness, and appropriateness of AI responses.
β’ Ability to articulate reasoning effectively to both technical and non-technical audiences.
β’ Meticulous attention to detail, including subtle errors, inconsistencies, and omissions.
β’ Consistent adherence to guidelines and quality standards.
β’ Flexibility to adapt to various projects, task types, and clients.
β’ Preferred: experience in adversarial machine learning, encompassing jailbreak datasets, prompt injection, RLHF/DPO attacks, and model extraction.
β’ Preferred: background in cybersecurity, including penetration testing, exploit development, and reverse engineering.
β’ Preferred: experience in socio-technical risk assessment, including probing for harassment/disinformation, abuse analysis, and conversational AI evaluation.
β’ Preferred: background in psychology, acting, or writing for innovative adversarial thinking.
β’ Must operate as an independent contractor.
β’ Must not require H1-B or STEM OPT sponsorship.
β’ Fully remote work environment.
β’ Flexible scheduling to meet personal needs.
β’ Weekly payments processed through Stripe or Wise.
β’ Competitive compensation.
β’ Access to wellness resources and clear guidelines for projects with higher sensitivity.
β’ Reasonable accommodations available upon request.
β’ Project durations may vary based on requirements and performance.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Chance to contribute to making AI systems more robust, safe, and trustworthy.
β’ Referral bonuses of up to $90 for each successful referral, subject to certain limits.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.