AI Safety Expert, English, Finnish

atMercorRemoteUS flagUnited StatesFreelanceArtificial IntelligenceMid-levelSenior$48 – $62/hour

Posted Sep 7

This is a fully remote position, open to applicants in United States.

πŸ“‹ Description

β€’ Conduct red-team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.

β€’ Create human data by annotating failures, categorizing vulnerabilities, and identifying systemic risks.

β€’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing practices.

β€’ Generate reproducible reports, datasets, and attack scenarios for clients.

β€’ Evaluate AI outputs related to sensitive subjects like bias, misinformation, or harmful behaviors.

β€’ Detect vulnerabilities that automated tests may overlook.

β€’ Broaden evaluation coverage and minimize unexpected issues in production.

β€’ Collaborate with top researchers to contribute to the training and enhancement of cutting-edge AI systems.


⛳️ Requirements

β€’ Native or fluent proficiency in English and Finnish.

β€’ Previous experience in red teaming related to AI adversarial work, cybersecurity, or socio-technical probing.

β€’ Capability to challenge systems adversarially and push them to their limits.

β€’ Proficiency in using frameworks or benchmarks for structured testing.

β€’ Ability to clearly communicate risks to both technical and non-technical stakeholders.

β€’ Flexibility to adapt across various projects and clients.

β€’ Preferred areas of expertise include adversarial ML, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation analysis, abuse investigation, conversational AI testing, psychology, acting, or creative adversarial writing.

β€’ Must operate as an independent contractor.

β€’ H1-B and STEM OPT candidates are not eligible.


🏝️ Benefits

β€’ Fully remote position.

β€’ Flexible work schedule; tasks can be completed at your convenience.

β€’ Weekly compensation through Stripe or Wise.

β€’ Opportunity to gain experience in human data-driven AI red teaming.

β€’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.

β€’ Competitive pay.

β€’ Collaboration opportunities with leading researchers.

β€’ Reasonable accommodations available upon request.

β€’ Referral bonuses of up to $250 for each successful referral.

People also viewed

WON.ai19 hours ago

AI Strategist

AR flagArgentina OnlyFreelanceArtificial Intelligence
ApplyView job
The College Board20 hours ago

Director, AI Assisted Solutions

US flagUnited States OnlyFull-timeArtificial Intelligence$88k – $135k/year
ApplyView job
Mercor21 hours ago

AI Safety Experts – English, Assamese

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor21 hours ago

AI Safety Experts – English, Marathi

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Mercor21 hours ago

AI Safety Expert – English, Bengali

US flagUnited States OnlyFreelanceArtificial Intelligence$16 – $22/hour
ApplyView job
Gartner21 hours ago

Director, Analyst – AI Technology Economics

GB flagUnited Kingdom OnlyFull-timeArtificial Intelligence
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers