
AI Safety Expert, English β Danish
Posted 7 hours ago

Posted 7 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents through methods such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human-generated data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure uniform testing practices.
β’ Generate reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs related to sensitive subjects, including bias, misinformation, and harmful behaviors.
β’ Contribute to the expansion of evaluation coverage and mitigate production surprises.
β’ Collaborate with top researchers and engage in projects aimed at training and improving AI systems.
β’ Native or fluent proficiency in both English and Danish.
β’ Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical probing.
β’ Capability to adversarially probe systems and push them to their limits.
β’ Familiarity with frameworks or benchmarks for systematic testing.
β’ Ability to clearly communicate risks to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects and client needs.
β’ Engagement as an independent contractor.
β’ Must not require H1-B or STEM OPT sponsorship.
β’ Preferred: experience in adversarial machine learning, including jailbreak datasets, prompt injections, RLHF/DPO attacks, or model extraction.
β’ Preferred: background in cybersecurity, encompassing penetration testing, exploit development, or reverse engineering.
β’ Preferred: experience in socio-technical risk assessment, including harassment/disinformation investigations, abuse analysis, or conversational AI evaluation.
β’ Preferred: creative probing experience in psychology, acting, or unconventional adversarial writing.
β’ Fully remote position that allows you to work on your own schedule.
β’ Weekly payments through Stripe or Wise based on services performed.
β’ Participation in higher-sensitivity projects is optional.
β’ Clear guidelines and wellness resources are provided for higher-sensitivity projects.
β’ Competitive compensation.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Referral bonuses of up to $250 for each successful referral.
β’ Reasonable accommodations available upon request.
anni.care
InnoData
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.