
AI Safety Expert β English, Portuguese
Posted Sep 4

Posted Sep 4
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Document failures, categorize vulnerabilities, and identify systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing.
β’ Create reproducible reports, datasets, and attack scenarios.
β’ Investigate AI outputs for bias, misinformation, and harmful behaviors.
β’ Assist in broadening evaluation coverage and enhancing customer AI systems.
β’ Native proficiency in English and Portuguese (global, not including Brazilian Portuguese).
β’ Previous experience in red teaming within AI adversarial fields, cybersecurity, or socio-technical probing.
β’ Capability to adversarially examine AI systems and stress them to their limits.
β’ Familiarity with frameworks or benchmarks for structured testing.
β’ Proficiency in clearly communicating risks to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects and clients.
β’ Independent contractor status is required.
β’ Must not hold H-1B or STEM OPT status.
β’ Preferred expertise includes adversarial machine learning, cybersecurity, socio-technical risk, or innovative probing techniques.
β’ Fully remote position.
β’ Flexible scheduling options.
β’ Weekly payments through Stripe or Wise.
β’ Participation in higher-sensitivity projects is optional.
β’ Access to clear content guidelines and wellness resources.
β’ Reasonable accommodations available upon request.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Collaboration with top researchers in the field.
β’ Referral bonuses of up to $180 for each successful referral.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.