
AI Safety Expert β English, Punjabi
Posted 12 hours ago

Posted 12 hours ago
This is a fully remote position, open to applicants in United States.
β’ Conduct red-team assessments on conversational AI models and agents.
β’ Investigate models for vulnerabilities such as jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Create human data by annotating failures and categorizing vulnerabilities.
β’ Identify systemic risks utilizing structured taxonomies, benchmarks, and playbooks.
β’ Develop reproducible reports, datasets, and attack scenarios for clients.
β’ Evaluate AI outputs concerning sensitive subjects like bias, misinformation, and harmful behaviors.
β’ Collaborate with top researchers and aid in the training and enhancement of cutting-edge AI systems.
β’ Native or fluent proficiency in English and Punjabi.
β’ Previous experience in red teaming within AI adversarial contexts, cybersecurity, or socio-technical investigations.
β’ Capability to test AI models through jailbreaks, prompt injections, misuse scenarios, bias exploitation, and multi-turn manipulations.
β’ Skill in annotating failures, classifying vulnerabilities, and identifying systemic risks.
β’ Familiarity with taxonomies, benchmarks, and playbooks.
β’ Competence in generating reproducible reports, datasets, and attack scenarios.
β’ Ability to clearly communicate risks to both technical and non-technical stakeholders.
β’ Flexibility to adapt across various projects and clientele.
β’ Preferred areas of expertise: adversarial machine learning, cybersecurity, socio-technical risk, or innovative probing techniques.
β’ Must be eligible to work independently as a contractor; H1-B and STEM OPT candidates are not supported.
β’ Fully remote work environment.
β’ Flexible schedule allowing you to work at your convenience.
β’ Weekly payments via Stripe or Wise based on services provided.
β’ Opportunity to gain experience in human data-driven AI red teaming.
β’ Direct involvement in enhancing the robustness, safety, and trustworthiness of AI systems.
β’ Collaborate with leading researchers in the field.
β’ Access to wellness resources and clear guidelines for higher-sensitivity projects.
β’ Reasonable accommodations available upon request.
β’ Referral bonuses of up to $90 for each successful referral.
Mercor
Riva Scientific
Mercor
Mercor
Get handpicked remote jobs straight to your inbox weekly.