
AI Safety Expert β English, Telugu
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments of conversational AI models and agents via jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation.
β’ Create human data by annotating failures, identifying vulnerabilities, and highlighting systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to ensure consistent testing processes.
β’ Develop reproducible reports, datasets, and attack scenarios for clients.
β’ Investigate AI outputs for biases, misinformation, and harmful behaviors.
β’ Identify vulnerabilities that automated tests might overlook.
β’ Broaden evaluation coverage and minimize unexpected issues in production.
β’ Enhance customer AI systems through adversarial testing.
β’ Fluent/native proficiency in English and Telugu is essential.
β’ Strong discernment regarding language and content, including evaluating the accuracy, completeness, and appropriateness of AI responses.
β’ Capability to articulate reasoning clearly to both technical and non-technical audiences.
β’ Meticulous attention to detail concerning subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines and quality standards.
β’ Flexibility across various projects, task types, and clients.
β’ Status as an independent contractor is required.
β’ Candidates must not be on an H1-B visa or STEM OPT.
β’ Preferred expertise includes adversarial machine learning, cybersecurity, socio-technical risk, creative probing, penetration testing, exploit development, reverse engineering, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, harassment/disinformation probing, abuse analysis, conversational AI testing, psychology, acting, or writing.
β’ Fully remote position that allows you to work on your own schedule.
β’ Weekly payments through Stripe or Wise based on services provided.
β’ Projects may be extended, shortened, or wrapped up early based on requirements and performance.
β’ Participation in higher-sensitivity projects is voluntary.
β’ Access to clear content guidelines and wellness resources for higher-sensitivity projects.
β’ Reasonable accommodations available upon request.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Chance to collaborate with leading researchers in the field.
β’ Referral payments of up to $90 for each successful referral.
Mercor
Mercor
Rove Concepts
Rove Concepts
Get handpicked remote jobs straight to your inbox weekly.