
AI Safety Expert β English, Tamil
Posted 6 days ago

Posted 6 days ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments of conversational AI models and agents through techniques such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by annotating failures, classifying vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs related to sensitive subjects like bias, misinformation, and harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Broaden evaluation scope and minimize unexpected issues in production.
β’ Collaborate with top researchers to aid in the training and enhancement of cutting-edge AI systems.
β’ Proficient in both English and Tamil languages.
β’ Strong discernment regarding language and content quality.
β’ Capability to evaluate whether AI responses are accurate, comprehensive, and suitable.
β’ Skill in identifying subtle errors, inconsistencies, and gaps in responses.
β’ Consistent adherence to guidelines and quality standards.
β’ Ability to clearly articulate reasoning to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects, task types, and clientele.
β’ Preferred experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Preferred background in cybersecurity, encompassing penetration testing, exploit development, or reverse engineering.
β’ Preferred experience in socio-technical risks, including harassment/disinformation probing, abuse analysis, or conversational AI testing.
β’ Preferred creative probing experience in fields such as psychology, acting, or writing.
β’ Must be eligible to work without H1-B or STEM OPT sponsorship.
β’ Engagement as an independent contractor is required.
β’ Fully remote working environment.
β’ Flexibility to set your own schedule.
β’ Weekly payments through Stripe or Wise.
β’ Competitive compensation.
β’ Participation in higher-sensitivity projects is optional.
β’ Clear guidelines and wellness resources available for managing sensitive content.
β’ Reasonable accommodations provided upon request.
β’ Referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.