AI Safety Expert, English, Dutch

atMercorRemoteUS flagUnited StatesFreelanceArtificial IntelligenceMid-levelSenior$48 – $62/hour

Posted 7 hours ago

This is a fully remote position, open to applicants in United States.

📋 Description

• Conduct red-team assessments on conversational AI models and agents by utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.

• Create human data through the annotation of failures, classification of vulnerabilities, and identification of systemic risks.

• Implement taxonomies, benchmarks, and playbooks to ensure consistent testing procedures.

• Generate reproducible reports, datasets, and attack case studies for clients.

• Analyze AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.

• Detect vulnerabilities that automated testing may overlook.

• Broaden evaluation coverage to minimize unexpected issues during production.

• Enhance customer AI systems through adversarial testing methodologies.


⛳️ Requirements

• Proficient/native proficiency in both English and Dutch.

• Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical investigations.

• Capability to challenge AI systems adversarially and push them to their limits.

• Familiarity with frameworks, taxonomies, benchmarks, or playbooks for structured testing approaches.

• Ability to articulate risks clearly to both technical and non-technical audiences.

• Flexibility to adapt across various projects and clientele.

• Status as an independent contractor.

• Must not require H1-B or STEM OPT sponsorship.

• Preferred specialties: adversarial machine learning, cybersecurity, socio-technical risk, or innovative probing techniques.

• Adversarial ML experience may involve working with jailbreak datasets, prompt injections, reinforcement learning from human feedback (RLHF)/differentiable programming optimization (DPO) attacks, or model extraction.

• Cybersecurity experience may include penetration testing, exploit development, or reverse engineering.

• Socio-technical risk experience may encompass harassment/disinformation probing, abuse analysis, or testing of conversational AI systems.

• Creative probing experience may involve psychology, acting, or writing skills.


🏝️ Benefits

• Fully remote position.

• Flexible working hours; tasks can be completed according to your own schedule.

• Weekly payments processed through Stripe or Wise.

• Project durations may be extended, shortened, or concluded early based on requirements and performance.

• Access to wellness resources and clear guidelines for projects with higher sensitivity.

• Opportunity to gain experience in human data-driven AI red teaming.

• Play a direct role in enhancing the robustness, safety, and trustworthiness of AI systems.

• Competitive compensation offered.

• Referral bonuses of up to $250 available for each successful referral, subject to certain limits.

• Reasonable accommodations can be requested as needed.

People also viewed

NICE5 hours ago

AI Solution Strategist

JP flagJapan OnlyFull-timeArtificial Intelligence
ApplyView job
BPCS, Comprehensive marketing solutions, ltd.7 hours ago

AI Response Labeler – Annotator

Latin AmericaFull-timeArtificial Intelligence$2,200 – $2,500/month
ApplyView job
Tech Mahindra8 hours ago

Conversational AI Designer – Amazon Connect, Lex NLU

US flagColorado OnlyFull-timeArtificial Intelligence$160k – $195k/year
ApplyView job
PwC10 hours ago

Senior Associate – Microsoft D365 ERP (F&O) AI/Copilot Functional Consultant

US flagArizona, +29 more statesFull-timeArtificial Intelligence$77k – $202k/year
ApplyView job
Creatio11 hours ago

AI Associate

PL flagPoland OnlyFull-timeArtificial Intelligence
ApplyView job
Veracity Consulting, Inc.12 hours ago

Technical Consultant – ServiceDesk, Now Assist, AI Control Tower

US flagUnited States OnlyFull-timeArtificial Intelligence
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers