
AI Safety Expert, English, Portuguese
Posted Aug 9

Posted Aug 9
This is a fully remote position, open to applicants in United States.
• Conduct red team assessments on conversational AI models and agents by utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
• Create human data by annotating failures, classifying vulnerabilities, and identifying systemic risks.
• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing methodologies.
• Generate reproducible reports, datasets, and attack scenarios for clients.
• Examine AI outputs related to sensitive issues like bias, misinformation, and harmful behaviors.
• Identify vulnerabilities overlooked by automated testing.
• Enhance evaluation coverage and minimize unexpected outcomes in production.
• Assist clients in making AI systems more robust, secure, and reliable.
• Collaborate with top researchers on initiatives focused on training and improving advanced AI systems.
• Native proficiency in English and Portuguese (global, excluding Brazilian Portuguese).
• Previous experience in red teaming related to AI adversarial tactics, cybersecurity, or socio-technical probing.
• Capability to challenge AI systems adversarially and push them to their limits.
• Familiarity with frameworks or benchmarks for systematic testing.
• Proficient in articulating risks clearly to both technical and non-technical audiences.
• Flexibility to adapt across various projects and client needs.
• Engagement as an independent contractor.
• Availability to undertake text-based tasks on a self-determined schedule.
• H1-B and STEM OPT candidates are not eligible.
• Preferred expertise includes adversarial machine learning, cybersecurity, socio-technical risk, and innovative probing techniques.
• Fully remote position.
• Flexible scheduling according to personal preferences.
• Weekly compensation through Stripe or Wise based on services provided.
• Access to wellness resources and clear protocols for projects that require heightened sensitivity.
• Competitive remuneration.
• Opportunity to collaborate with leading researchers in the field.
• Chance to gain experience in human data-driven AI red teaming.
• Reasonable accommodations available upon request.
• Referral bonuses of up to $180 for each successful referral.
CVS Health
One Impression
Volga Partners
Mercor
Get handpicked remote jobs straight to your inbox weekly.