
Staff Research Engineer – RRI Harness
Posted 23 hours ago

Posted 23 hours ago
This is a fully remote position, open to applicants in United States.
• Develop and enhance the Python multi-agent runtime, focusing on orchestration, tool contracts, context and memory management, compaction, and model adapters across various LLM providers.
• Design and conduct experiments related to agent prompting, tool design, context and memory strategies, and model selection.
• Implement changes that significantly improve RRI results.
• Provide a fixed evaluation set, a scorecard with measured noise bands, evaluations for each change, and component ablations.
• Engineer resilience against GPU failures, node replacements, rolling upgrades, provider errors, and credit limits.
• Create event streams, perform cross-session analyses, manage per-agent cost and token accounting, and develop MCP administration tools.
• Identify and minimize agent idling, looping, and excessive token utilization.
• Design interfaces for the Go control plane and GPU proxy owners.
• Collaborate with applied scientists and forward-deployed engineers to troubleshoot and resolve client issues in production.
• Take ownership of the core runtime for Sequen's autonomous research engine.
• Over 7 years of experience in developing production software.
• Proficient in Python.
• Experience with distributed or long-running systems.
• Familiarity with LLMs in agentic loops, including tool calling, prompt and context management, streaming, and retries.
• Capable of debugging across processes, containers, and network boundaries.
• Experience with Kubernetes and GPU nodes.
• Ability to interpret a training script, comprehend a ranking metric, and differentiate actual regressions from noise.
• Applied research experience with ML or LLM experiments, including baselines and ablations.
• Experience with evaluation infrastructure, test harnesses, benchmarks, experiment tracking, or ML CI.
• Basic knowledge of Go.
• Familiarity with on-prem, BYOC, or air-gapped environments and their security constraints.
• Experience with large-scale PyTorch training, GPU scheduling, or ML platforms.
• Contributions to agent frameworks, evaluation harnesses, or ML tooling are considered a plus.
• Unlimited paid time off.
• Flexible hybrid/remote work configurations.
• A highly collaborative, world-class engineering culture.
• Equity.
Sequen
Burq
Citrine Informatics
Mortenson
Get handpicked remote jobs straight to your inbox weekly.