
AI Safety Expert, English β Odia
Posted Sep 6

Posted Sep 6
This is a fully remote position, open to applicants in United States.
β’ Conduct red team activities on conversational AI models and agents
β’ Investigate models for vulnerabilities such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation
β’ Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks
β’ Utilize taxonomies, benchmarks, and playbooks to maintain consistency in testing
β’ Generate reproducible reports, datasets, and attack cases for client use
β’ Assess AI outputs related to sensitive topics like bias, misinformation, and harmful behaviors
β’ Enhance evaluation coverage and reveal vulnerabilities that automated tests may overlook
β’ Develop artifacts that enhance the robustness of customer AI systems
β’ Required fluent/native proficiency in both English and Odia
β’ Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical probing
β’ Capability to probe AI models through techniques such as jailbreaks, prompt injections, misuse cases, bias exploitation, and multi-turn manipulation
β’ Skills in annotating failures, classifying vulnerabilities, and identifying systemic risks
β’ Ability to adhere to taxonomies, benchmarks, and playbooks
β’ Proficiency in producing reproducible reports, datasets, and attack scenarios
β’ Capability to articulate risks clearly to both technical and non-technical stakeholders
β’ Flexibility across various projects and client needs
β’ Preferred specialties include adversarial machine learning, cybersecurity, socio-technical risk, and innovative probing
β’ Engagement as an independent contractor
β’ H1-B and STEM OPT candidates are not eligible
β’ Fully remote position
β’ Flexible working hours; complete tasks on your own schedule
β’ Receive weekly payments through Stripe or Wise
β’ Chance to gain experience in human data-driven AI red teaming at the forefront of safety
β’ Direct contribution to enhancing the robustness, safety, and trustworthiness of AI systems
β’ Collaborate with leading researchers in the field
β’ Competitive compensation
β’ Participation in higher-sensitivity projects is optional
β’ Clear guidelines and wellness resources available for higher-sensitivity projects
β’ Reasonable accommodations provided upon request
β’ Referral bonuses of up to $90 for each successful referral, with no limit on the number of referrals
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.