
AI Safety Expert, English & Assamese
Posted 6 hours ago

Posted 6 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing protocols.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs related to sensitive subjects such as bias, misinformation, or harmful behaviors.
β’ Detect vulnerabilities that automated tests may overlook.
β’ Broaden evaluation coverage and minimize unexpected production issues.
β’ Enhance customer AI systems through adversarial testing strategies.
β’ Proficiency in both English and Assamese is required.
β’ Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical probing.
β’ Capability to adversarially probe AI systems and push them to their limits.
β’ Familiarity with frameworks or benchmarks for structured testing methodologies.
β’ Ability to communicate risks effectively to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects and client needs.
β’ Must hold independent contractor status.
β’ Unable to provide support for H1-B or STEM OPT candidates.
β’ Nice-to-have: experience in adversarial machine learning, including jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Nice-to-have: background in cybersecurity, including penetration testing, exploit development, or reverse engineering.
β’ Nice-to-have: experience with socio-technical risks, including harassment/disinformation probing, abuse analysis, or testing of conversational AI.
β’ Nice-to-have: creative probing skills in psychology, acting, or writing.
β’ Flexible working hours; tasks can be completed according to your own schedule.
β’ Weekly payments via Stripe or Wise based on services provided.
β’ Fully remote work environment.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Collaboration with top researchers in the field.
β’ Competitive compensation.
β’ Participation in higher-sensitivity projects is optional.
β’ Clear guidelines and wellness resources available for projects involving sensitive content.
β’ Referral bonuses of up to $90 for each successful referral.
XenoPatch GmbH
Mercor
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.