
AI Safety Experts β English, Marathi
Posted 1 day ago

Posted 1 day ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Generate human data by annotating failures, classifying vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing.
β’ Create reproducible reports, datasets, and attack scenarios for clients.
β’ Review AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Detect vulnerabilities that automated testing may overlook.
β’ Increase evaluation coverage and minimize unexpected production issues.
β’ Assist in training and advancing frontier AI systems for Mercor's AI lab and enterprise clientele.
β’ Fluent/native proficiency in English and Marathi.
β’ Strong discernment regarding language and content.
β’ Capability to evaluate whether AI responses are accurate, comprehensive, and suitable, along with the ability to articulate the reasoning.
β’ Meticulous attention to subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines and quality standards.
β’ Ability to clearly communicate reasoning to both technical and non-technical audiences.
β’ Flexibility across various projects, task types, and clients.
β’ Status as an independent contractor.
β’ H1-B and STEM OPT candidates are not eligible.
β’ Nice-to-have: experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, and model extraction.
β’ Nice-to-have: background in cybersecurity, including penetration testing, exploit development, and reverse engineering.
β’ Nice-to-have: experience with socio-technical risks, such as harassment/disinformation probing, abuse analysis, and conversational AI testing.
β’ Nice-to-have: experience in psychology, acting, or writing to foster unconventional adversarial thinking.
β’ Fully remote position.
β’ Flexible work schedule allowing you to manage your own time.
β’ Weekly payments through Stripe or Wise.
β’ Project durations may vary depending on needs and performance.
β’ Reasonable accommodations available upon request.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Competitive compensation.
β’ Collaboration with leading researchers.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.