Remotery

Head of AI Safety

Posted Aug 7

This is a fully remote position, open to applicants in Colorado, +11 more states.

📋 Description

• Oversee and ensure the quality of Moonshot's applied AI safety initiatives focusing on violence, extremism, CSEA, abuse and grooming, mental health and crisis, as well as child and teen risk factors.

• Provide guidance to frontier AI companies on enhancing the safety of models, products, policies, and interventions.

• Convert specialized insights into practical advice for model safety, policy-making, product development, research, and engineering teams.

• Establish methodological strategies and structured, testable evaluation frameworks.

• Lead and engage in red teaming and adversarial evaluations utilizing test scenarios, model responses, scoring criteria, safety policies, and evaluation outcomes.

• Detect safety deficiencies and create recommendations for model behavior and user protections.

• Maintain meticulous documentation for technical, governmental, and foundation stakeholders.

• Ensure compliance with ethical, legal, contractual, data protection, and regulatory requirements.

• Manage risks related to operations, reputation, delivery, and partnerships.

• Act as Moonshot's primary point of contact for applied AI safety with AI companies, governmental bodies, regulators, and ecosystem partners.

• Cultivate relationships with technical teams, governmental agencies, foundations, academics, researchers, civil society organizations, and practitioners.

• Represent Moonshot in various meetings, briefings, workshops, and sector engagements.

• Lead, mentor, and manage the AI safety team; facilitate workforce planning, performance assessments, and professional growth.

• Collaborate with operations, finance, research, and technical teams.

• Develop the AI safety portfolio, strategic initiatives, partnerships, and funding opportunities.

• Spearhead proposal development, scoping, renewals, communications, publications, briefings, and thought leadership endeavors.

• Create repeatable methodologies and service offerings while ensuring quality and rigor in delivery.

• Oversee project planning, staffing, budgeting, forecasting, and delivery timelines.


⛳️ Requirements

• Proven experience in trust & safety, online harms, violence prevention, safeguarding, public health, or a closely related discipline.

• Capacity to apply relevant harm-prevention knowledge to AI systems effectively.

• Genuine curiosity about AI and the ability to rapidly develop technical fluency.

• Experience in designing research, evaluation frameworks, or interventions addressing violent extremism, CSEA, self-harm and crisis, or targeted violence.

• Proficiency in managing projects, teams, budgets, partners, and client relationships.

• Strong skills in people management.

• Exceptional written communication skills and experience creating credible materials for government, foundation, or enterprise audiences.

• Resilience and comfort in handling highly sensitive or graphic content.

• Sound judgment and the ability to navigate ambiguity, competing priorities, and delicate stakeholder environments.

• Willingness to travel and work beyond standard hours when necessary.

• Trustworthiness, discretion, diplomacy, and readiness to undergo relevant security clearance procedures.

• Experience in supporting business development, grant funding, or procurement activities.

• A strong commitment to Moonshot's mission.

• Eligibility to work in the US.

• Capability to pass relevant security clearance processes as per client requirements.

• Desirable: experience in model safety, red teaming, or adversarial evaluation of LLMs or other AI systems.

• Desirable: knowledge of LLM architecture, safety tooling, or trust & safety policies.

• Desirable: experience in child safety evaluation, teen-safety product development, or grooming and CSEA detection.

• Desirable: experience engaging with government or regulatory bodies.

• Desirable: background in designing intervention or diversion programs.

• Desirable: academic or applied expertise in radicalization studies, forensic psychology, or violence risk assessment.

• Desirable: familiarity with taxonomy or classifier development and testing data.


🏝️ Benefits

• 15 days of paid vacation leave, in addition to Federal holidays and an extra paid leave day for Native American Heritage Day.

• Flexible public holiday policy, allowing the option to work federal holidays in exchange for time off at another occasion.

• Comprehensive private healthcare package that includes coverage for partners and children.

• Dental & Vision Insurance.

• Life & Disability Insurance.

• 24/7 access to complimentary counseling through our Employee Assistance Program.

• 3% matched 401k contributions.

• 401(k) Roth Contributions.

• Generous maternity and paternity leave: 26 weeks of paid maternity leave and 8 weeks of paid paternity leave.

• All permanent employees are eligible for share options upon starting employment.

People also viewed

Progressive Leasing9 hours ago

AI Workforce Enablement Consultant – Contract

US flagArizona OnlyFull-timeArtificial Intelligence
ApplyView job
apna10 hours ago

AI Voice Data Collection – Hindi

IN flagIndia OnlyFreelanceArtificial Intelligence₹500/hour
ApplyView job
apna10 hours ago

AI Voice Data Collection – Odia

IN flagIndia OnlyFreelanceArtificial Intelligence₹500/hour
ApplyView job
Texas Research International10 hours ago

Junior Software and Systems Specialist – AI and Automation

US flagUnited States OnlyFull-timeArtificial Intelligence
ApplyView job
Welo Global11 hours ago

Generative AI Analyst, English

GB flagUnited Kingdom OnlyFreelanceArtificial Intelligence$19/hour
ApplyView job
Welo Global11 hours ago

Generative AI Analyst – French

CA flagCanada OnlyFreelanceArtificial Intelligence$20/hour
ApplyView job

Never miss a great job!

Get handpicked remote jobs straight to your inbox weekly.

Trusted by 7,400+ designers