
Head of Cyber Safety
Posted 2 days ago

Posted 2 days ago
This is a fully remote position, open to applicants in United States.
β’ Lead and design adversarial evaluations of advanced LLMs focusing on offensive cyber capabilities, such as vulnerability identification, exploit creation, malware production, privilege escalation, social engineering, persistence, and autonomous cyber operations across text, agentic, and multimodal systems.
β’ Collaborate with machine learning engineers to convert cybersecurity knowledge into scalable benchmarks, classifiers, guardrails, automated detection frameworks, and evaluation infrastructure.
β’ Develop and oversee Gray Swanβs catastrophic cyber harm taxonomy, adapting cyber evaluation frameworks as the capabilities of frontier models progress.
β’ Generate technical risk assessments and provide actionable recommendations for frontier AI laboratories, enterprise clients, and internal stakeholders.
β’ Build, guide, and lead a team of cybersecurity subject matter experts while establishing scalable evaluation processes, quality benchmarks, and technical infrastructure.
β’ Serve as Gray Swanβs cybersecurity authority, engaging with frontier AI laboratories, security researchers, government entities, and communities involved in AI safety and cybersecurity.
β’ Extensive technical knowledge in offensive cybersecurity, vulnerability research, exploit development, penetration testing, malware analysis, reverse engineering, or a closely related field through industry experience, research, or equivalent qualifications.
β’ Considerable experience in evaluating advanced cyber threats, offensive tools, or AI-driven cyber capabilities, particularly within critical infrastructure sectors.
β’ Practical experience in conducting adversarial evaluations, AI red-teaming, LLM security research, or creating evaluation datasets for advanced AI systems.
β’ Comfortable working at the intersection of cybersecurity research, AI safety, and machine learning engineering.
β’ Proven ability to navigate highly ambiguous and fast-paced environments while formulating strategies and developing new capabilities.
β’ Experience in developing machine learning models, AI security classifiers, or automated cyber detection systems is an added advantage.
β’ Background in red-teaming frontier language models, jailbreaking, prompt injection research, or evaluations of agentic AI is a plus.
β’ Familiarity with frontier AI laboratories, national security agencies, or leadership roles in cybersecurity research teams is a plus.
β’ A background in threat intelligence, autonomous cyber operations, AI agent security, or AI governance is advantageous.
β’ Strong software engineering skills in Python, Go, Rust, or other systems programming languages are considered a bonus.
β’ Must possess legal authorization to work in the United States.
β’ Performance-based bonus
β’ Meaningful equity package
β’ 401k with up to 4% matching
β’ 28 days of annual leave (including vacation and holidays)
β’ Health, dental, and vision insurance coverage
β’ Catered lunches (available at the Pittsburgh office)
β’ Flexible work arrangements
β’ Visa sponsorship available for exceptional candidates
Behavioral Health Works, Inc.
Sodexo
Sodexo
EVERSANA
Get handpicked remote jobs straight to your inbox weekly.