
AI Safety Expert β English, Punjabi
Posted 9 hours ago

Posted 9 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents by exploiting jailbreaks, prompt injections, misuse scenarios, bias issues, and multi-turn manipulations.
β’ Create high-quality human data through the annotation of failures, classification of vulnerabilities, and identification of systemic risks.
β’ Utilize taxonomies, benchmarks, and playbooks to maintain consistency in testing procedures.
β’ Generate reproducible reports, datasets, and attack scenarios that clients can implement.
β’ Evaluate AI outputs related to sensitive subjects such as bias, misinformation, and harmful behaviors.
β’ Identify vulnerabilities that automated testing may overlook.
β’ Broaden evaluation scope by testing additional scenarios and minimizing production surprises.
β’ Native proficiency in both English and Punjabi is mandatory.
β’ Strong judgment regarding language and content, including the ability to assess the accuracy, completeness, and appropriateness of AI-generated responses.
β’ Capability to articulate reasoning clearly to both technical and non-technical audiences.
β’ Keen attention to detail for identifying subtle errors, inconsistencies, and gaps.
β’ Consistent adherence to guidelines and quality standards.
β’ Flexibility to adapt to various projects, task types, and client needs.
β’ Engagement as an independent contractor.
β’ Candidates must not be on an H1-B visa or STEM OPT.
β’ Preferred: experience in adversarial machine learning, including knowledge of jailbreak datasets, prompt injection, RLHF/DPO attacks, or model extraction.
β’ Preferred: background in cybersecurity, including penetration testing, exploit development, or reverse engineering.
β’ Preferred: experience with socio-technical risks, such as harassment/disinformation probing, abuse analysis, or conversational AI testing.
β’ Preferred: creative probing experience in psychology, acting, or writing.
β’ Fully remote position that allows for flexible scheduling.
β’ Weekly compensation through Stripe or Wise based on services performed.
β’ Participation in higher-sensitivity projects is optional.
β’ Access to clear guidelines and wellness resources for sensitive-content projects.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Competitive compensation.
β’ Collaboration with leading researchers in the field.
β’ Reasonable accommodations available upon request.
Genesys
Mercor
Genesys
Get handpicked remote jobs straight to your inbox weekly.