
Staff Research Engineer – Recursive Self-Improvement Lead
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in United States.
• Develop the research roadmap for recursive self-improvement and determine which projects to pursue.
• Create and evaluate exploration policies, manage resource allocation among parallel agents, and enable agents to modify their own instructions and strategies.
• Transform recorded research sessions into training and evaluation signals.
• Establish criteria for improved research outcomes and collaborate with the RRI team to develop evaluations.
• Ensure that knowledge acquired in one session enhances future sessions across various tasks and clients.
• Monitor literature on RSI, autonomous research, and agentic LLMs, converting promising concepts into quantifiable experiments.
• Collaborate with applied scientists utilizing RRI to tackle real client challenges and identify limitations in agents.
• Write code and successfully integrate research findings into systems utilized by clients in production.
• Contribute to the growth of the Recursive Self-Improvement team.
• Over 7 years of experience in applied ML research or research engineering.
• Practical experience in at least one of the following areas: LLM agents, AutoML, meta-learning, recursive self-improvement, or a closely related field.
• Proven ability to transition research ideas from prototype stages to measurable enhancements in deployed products or production systems.
• Strong experimental methodology, including baselines, ablations, seeds, and confidence intervals.
• Proficiency in production-quality Python and experience in conducting end-to-end experiments.
• Knowledge of frontier model behavior in prolonged agentic loops, encompassing context limits, compaction, tool usage, failure modes, and associated costs.
• Capacity to take ownership of a research direction within an early-stage team.
• Experience in leading or mentoring a small research team.
• Peer-reviewed publications or significant open-source contributions in areas such as agents, RL, or AutoML may be beneficial.
• Familiarity with reinforcement learning or post-training methods for LLMs, including RLHF, reward modeling, or training agents using RL may be advantageous.
• Experience in developing evaluations or benchmarks for LLMs or agents may be beneficial.
• A background in search, recommendation, or learning-to-rank models may be advantageous.
• Familiarity with multi-agent systems or agent swarms may be beneficial.
• Unlimited paid time off.
• Flexible hybrid/remote work options.
• A highly collaborative, world-class engineering culture.
• Equity.
Sequen
Burq
Citrine Informatics
Mortenson
Get handpicked remote jobs straight to your inbox weekly.