
Research Lead – Pre-training Safety
Posted Sep 8

Posted Sep 8
This is a fully remote position, open to applicants in United States, +2 more locations.
• Spearhead and oversee FAR.AI’s pre-training safety research initiatives.
• Manage capability-control research aimed at eliminating harmful capabilities while retaining benign ones.
• Scale techniques such as Deep Ignorance for models exceeding 100B parameters and 1T tokens.
• Collaborate with the red team to rigorously test models and assess scaling to frontier systems.
• Investigate enhanced data filtering techniques utilizing data attribution, influence-based selection, and advanced classifiers.
• Apply gradient routing to identify dual-use capabilities within model components like MoE experts.
• Create training methodologies to actively eliminate harmful capabilities through unlearning and next-token prediction.
• Introduce synthetic data during pre-training or mid-training to shape model representations and behaviors.
• Build and guide the team, establishing its research objectives.
• Mentor Technical Staff members while remaining actively involved in coding and conducting experiments.
• Define a research agenda and a theory of change for mitigating catastrophic AI risks or enhancing beneficial outcomes.
• Lead innovative research projects that may have uncertain progress or success indicators.
• Disseminate findings via academic publications, blog entries, ML conferences, and briefings for policymakers.
• Contribute to FAR.AI’s intellectual atmosphere and research culture.
• Foster a research area through grantmaking and events.
• Link research outcomes to real-world applications through independent testing and advising government entities.
• Proven research experience in AI or another highly technical domain, such as computer science, mathematics, or physics.
• Extensive background in language-model pretraining, dataset construction, or controlled training experiments.
• Experience in developing large-scale pipelines for scoring, filtering, deduplicating, and sampling training corpora.
• Strong experimental judgment, encompassing safety-capability evaluations, distribution-shift analysis, and statistically sound model comparisons.
• Capability to construct and troubleshoot research systems directly, ranging from classifier fine-tuning to distributed training and evaluation.
• A defined research agenda with a theory of change, or a proven track record and research area that can evolve into an agenda.
• Experience in leading a team, mentoring graduate students, or supporting early-career researchers; informal leadership is also recognized.
• Proficiency in conveying novel methods and solutions to both technical and non-technical audiences.
• Significant involvement in machine learning research through prior research roles, employment, or sustained independent contributions; newcomers are not preferred.
• Preferred: Established publication record in AI safety.
• Preferred: Comfort with writing grant proposals and managing external collaborations.
• Availability for full-time work, 40 hours per week.
• Ability to work remotely in most countries or in-person in Berkeley, California, or Singapore.
• Competitive salary ranging from $290,000 to $450,000 per year, based on experience and location; exceptional candidates may be offered a higher salary.
• Coverage for work-related travel and equipment expenses.
• Catered lunch and dinner provided at Berkeley offices.
• Flexible options for remote or in-person work.
• Visa sponsorship available for in-person employees.
• Paid trial period of 3 to 5 days.
• Generous compute budgets on a managed cluster.
• Dedicated engineering team to support compute infrastructure and experiment scaling.
• Opportunities for collaboration with governments, leading AI organizations, and academic institutions.
• Professional development through research leadership, mentoring, publications, conferences, grantmaking, and events.
Mercor
CBT
Compass Experience Labs
Bird's Eye Medical
Get handpicked remote jobs straight to your inbox weekly.