
Head of AI Safety
Posted 22 hours ago

Posted 22 hours ago
This is a fully remote position, open to applicants in Ireland.
• Oversee the delivery, development, and expansion of Moonshot's AI Safety portfolio.
• Direct and ensure quality in applied AI safety initiatives encompassing violence, extremism, CSEA, abuse and grooming, mental health and crisis, as well as child and teen risk categories.
• Provide guidance to frontier AI companies on enhancing model, product, policy, and intervention safety.
• Convert subject-matter expertise into practical recommendations for model safety, policy, product, research, and engineering teams.
• Establish methodological approaches and create structured, testable evaluation frameworks.
• Actively lead and engage in red teaming and adversarial evaluation activities.
• Detect safety failures and formulate suggestions for improving model behavior and user protections.
• Ensure meticulous documentation while adhering to ethical, legal, contractual, data protection, and compliance standards.
• Manage risks related to operations, reputation, delivery, and partnerships.
• Serve as Moonshot's principal applied AI safety representative for AI companies, governments, regulators, and ecosystem partners.
• Cultivate relationships with technical teams, governments, foundations, regulators, academics, researchers, civil society organizations, and practitioners.
• Represent Moonshot in external meetings, briefings, workshops, and sector engagements.
• Lead, mentor, and manage the AI safety team.
• Assist in workforce planning, performance management, professional development, and team well-being.
• Collaborate with operations, finance, research, and technical teams.
• Expand the AI safety portfolio through strategic opportunities, partnerships, and funding sources.
• Spearhead proposal development, scoping, renewals, and business development efforts.
• Create repeatable methodologies, service offerings, and partnerships.
• Aid in communications, publications, briefings, and thought leadership initiatives.
• Supervise project planning, staffing, budgeting, forecasting, and delivery timelines.
• Proven experience in trust & safety, online harms, violence prevention, safeguarding, or public health, with the ability to adapt knowledge to AI systems.
• A strong curiosity about AI and the capacity to quickly develop technical fluency.
• Experience in designing research, evaluation frameworks, or interventions related to violent extremism, CSEA, self-harm and crisis, or targeted violence.
• Demonstrated ability to manage projects, teams, budgets, partners, and clients effectively.
• Strong skills in people management.
• Exceptional written communication skills tailored for government, foundation, or enterprise audiences.
• Resilience and comfort in handling highly sensitive or graphic content.
• Awareness of well-being practices when dealing with sensitive content.
• Sound judgment and the ability to navigate ambiguity, competing priorities, and sensitive stakeholder environments.
• Willingness to travel and work outside regular hours when necessary.
• Trustworthiness, discretion, diplomacy, and readiness to undergo relevant security clearance procedures.
• Experience in supporting business development, grant funding, or procurement.
• Strong commitment to Moonshot's mission.
• Eligibility to work in the UK.
• Ability to pass necessary security clearance procedures.
• Direct experience in model safety, red teaming, or adversarial evaluation of LLMs or other AI systems (desirable).
• Understanding of LLM architecture, safety tooling, or trust & safety policy (desirable).
• Experience in child safety evaluation, teen-safety product work, or grooming and CSEA detection (desirable).
• Experience in government or regulatory engagement (desirable).
• Experience in designing intervention or diversion programs (desirable).
• Academic or applied background in radicalisation studies, forensic psychology, or violence risk assessment (desirable).
• Familiarity with taxonomy or classifier development and testing data workflows (desirable).
• 30 days of paid annual leave, excluding public holidays.
• Flexible public holiday policy allowing the option to work public holidays in exchange for a day off at another time.
• Private healthcare package that includes access to specialist mental health coverage for partners and children.
• Dental and vision insurance.
• Life insurance and income protection.
• Employee Assistance Programme providing access to mental health support.
• 26 weeks of paid maternity leave.
• 8 weeks of paid paternity leave.
• Share options granted to all permanent employees upon employment.
Mashreq
LILT AI
BCD Travel
BCD Travel
Get handpicked remote jobs straight to your inbox weekly.