
AI Safety Expert, English β Punjabi
Posted Sep 18

Posted Sep 18
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation techniques.
β’ Document failures, categorize vulnerabilities, and identify systemic risks.
β’ Consistently apply taxonomies, benchmarks, and playbooks.
β’ Generate reproducible reports, datasets, and attack scenarios.
β’ Evaluate AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Identify vulnerabilities that automated tests might overlook.
β’ Enhance evaluation coverage and minimize unexpected outcomes in production.
β’ Fortify customer AI systems through data-driven human red teaming.
β’ Native proficiency in English and Punjabi.
β’ Strong discernment regarding language use and content quality.
β’ Capability to evaluate if AI responses are accurate, complete, and suitable, along with the ability to articulate the reasoning behind assessments.
β’ Keen eye for detecting subtle errors, inconsistencies, and omissions.
β’ Consistent adherence to guidelines and quality standards.
β’ Ability to communicate reasoning effectively to both technical and non-technical audiences.
β’ Flexibility across various projects, task types, and clientele.
β’ Status as an independent contractor.
β’ Not requiring H1-B or STEM OPT sponsorship.
β’ Preferred expertise areas: adversarial machine learning, cybersecurity, socio-technical risk, or creative probing.
β’ Fully remote position.
β’ Flexible, self-directed work schedule.
β’ Weekly compensation through Stripe or Wise.
β’ Project durations may be adjusted or concluded early based on requirements and performance.
β’ Optional involvement in projects with higher sensitivity.
β’ Clear content guidelines and access to wellness resources.
β’ Reasonable accommodations available upon request.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Collaboration with top researchers in the field.
β’ Referral bonuses of up to $90 for successful referrals.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.