
AI Safety Experts, English β Marathi
Posted Sep 19

Posted Sep 19
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents through various methods such as jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation.
β’ Document failures, categorize vulnerabilities, and identify systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing.
β’ Generate reproducible reports, datasets, and attack scenarios.
β’ Evaluate AI outputs related to sensitive subjects, including bias, misinformation, and harmful behaviors.
β’ Assist clients in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Collaborate across diverse projects, task types, and clientele.
β’ Proficiency in both English and Marathi is mandatory.
β’ Strong discernment regarding language and content.
β’ Capability to evaluate the accuracy, completeness, and appropriateness of AI responses, with the ability to articulate reasoning.
β’ Keen attention to detail, able to spot subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines and quality standards.
β’ Ability to clearly communicate reasoning to both technical and non-technical audiences.
β’ Flexibility across various projects, task types, and clients.
β’ Must hold independent contractor status.
β’ H1-B and STEM OPT candidates are not eligible.
β’ Preferred expertise includes adversarial ML, cybersecurity, socio-technical risk, and creative probing.
β’ Fully remote work environment.
β’ Flexible working hours, allowing you to set your own schedule.
β’ Weekly compensation via Stripe or Wise.
β’ Competitive salary.
β’ Opportunity to work alongside leading researchers.
β’ Gain experience in human data-driven AI red teaming.
β’ Reasonable accommodations available upon request.
β’ Referral bonuses of up to $90 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.