
AI Safety Expert, English, Marathi
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to maintain consistency in testing.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs related to sensitive issues like bias, misinformation, or harmful behaviors.
β’ Identify vulnerabilities that automated tests may overlook.
β’ Broaden evaluation coverage and minimize unexpected issues in production.
β’ Assist clients in enhancing the safety, robustness, and reliability of AI systems.
β’ Required fluent/native proficiency in English and Marathi.
β’ Strong judgment regarding language and content; capability to evaluate the accuracy, completeness, and appropriateness of AI responses while providing clear explanations.
β’ Meticulous attention to minor errors, inconsistencies, and omissions.
β’ Consistent adherence to guidelines and quality standards.
β’ Ability to articulate reasoning effectively to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects, task types, and client needs.
β’ Must hold independent contractor status.
β’ Note: H1-B and STEM OPT candidates are not eligible.
β’ Preferred: Experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Preferred: Background in cybersecurity, covering penetration testing, exploit development, or reverse engineering.
β’ Preferred: Experience in socio-technical risk assessment, such as harassment/disinformation probing, abuse analysis, or conversational AI evaluation.
β’ Preferred: Background in psychology, acting, or writing to foster unconventional adversarial thinking.
β’ Fully remote position.
β’ Flexible working hours; manage your own schedule.
β’ Weekly payments via Stripe or Wise based on services provided.
β’ Project durations can be adjusted based on requirements and performance.
β’ Access to wellness resources and clear protocols for high-sensitivity projects.
β’ Reasonable accommodations available upon request.
β’ Referral program: earn up to $90 for each successful referral, with no cap on the number of referrals.
β’ Opportunity to collaborate with leading researchers.
β’ Gain experience in human data-driven AI red teaming at the forefront of safety.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.