
AI Safety Expert β English, Swedish
Posted Aug 24

Posted Aug 24
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents, utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.
β’ Generate reproducible reports, datasets, and attack scenarios for client use.
β’ Evaluate AI outputs related to sensitive issues such as bias, misinformation, and harmful behaviors.
β’ Broaden evaluation coverage and reveal vulnerabilities that automated tests may overlook.
β’ Engage in projects aimed at training and enhancing cutting-edge AI systems.
β’ Native or fluent proficiency in both English and Swedish.
β’ Previous experience in red teaming within the realms of AI adversarial work, cybersecurity, or socio-technical probing.
β’ Capability to probe systems adversarially and challenge them to their limits.
β’ Experience with frameworks or benchmarks for systematic testing methodologies.
β’ Proficiency in articulating risks clearly to both technical and non-technical audiences.
β’ Flexibility to adapt across various projects and clientele.
β’ Experience in adversarial machine learning, including but not limited to jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction (preferred).
β’ Background in cybersecurity specialties such as penetration testing, exploit development, or reverse engineering (preferred).
β’ Familiarity with socio-technical risk assessment, including harassment/disinformation probing, abuse analysis, or conversational AI evaluation (preferred).
β’ Creative probing experience in fields like psychology, acting, or writing (preferred).
β’ Candidates must not be on an H-1B or STEM OPT visa.
β’ Fully remote position.
β’ Flexible work schedule allowing you to set your own hours.
β’ Receive weekly payments via Stripe or Wise.
β’ Gain valuable experience in human data-driven AI red teaming.
β’ Play a direct role in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Competitive compensation.
β’ Collaborate with leading researchers in the field.
β’ Reasonable accommodations available upon request.
β’ Earn referral bonuses of up to $250 for each successful referral (conditions may apply).
Mercor
Mercor
Mercor
Quality Digital
Get handpicked remote jobs straight to your inbox weekly.