
AI Safety Expert β English, Telugu
Posted 5 days ago

Posted 5 days ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents utilizing jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulation techniques.
β’ Document failures, categorize vulnerabilities, and identify systemic risks.
β’ Implement taxonomies, benchmarks, and playbooks to ensure consistent testing processes.
β’ Generate reproducible reports, datasets, and attack case studies for clients.
β’ Analyze AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Collaborate on initiatives aimed at training and enhancing AI systems.
β’ Identify vulnerabilities that automated testing might overlook.
β’ Broaden evaluation coverage and minimize unexpected issues in production.
β’ Required fluent/native proficiency in both English and Telugu.
β’ Strong discernment regarding language and content, including assessing the accuracy, completeness, and appropriateness of AI responses.
β’ Ability to articulate reasoning clearly to both technical and non-technical audiences.
β’ Meticulous attention to subtle errors, inconsistencies, and omissions.
β’ Consistently adhere to guidelines and quality standards.
β’ Flexibility across various projects, task types, and clientele.
β’ Engagement as an independent contractor.
β’ Candidates must not hold H1-B or STEM OPT status.
β’ Preferred specialties include adversarial machine learning, jailbreak datasets, prompt injection, RLHF/DPO attacks, model extraction, penetration testing, exploit development, reverse engineering, harassment/disinformation investigation, abuse analysis, conversational AI testing, psychology, acting, and writing.
β’ Fully remote position that allows you to work on your own schedule.
β’ Weekly payments through Stripe or Wise based on services performed.
β’ Opportunity to gain experience in human data-driven AI red teaming at the cutting edge of safety.
β’ Play a direct role in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Participation in higher-sensitivity projects is optional.
β’ Clear guidelines and wellness resources available for higher-sensitivity projects.
β’ Reasonable accommodations provided upon request.
β’ Referral bonuses of up to $90 for each successful referral, with no specified limit on the number of referrals.
The College Board
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.