跳到正文
原文
HuggingFace Daily Papers(社區熱門論文)·· 2 天前AI 評分50

DAEDALUS:以自生成任務引導建立智能體記憶

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

AI 導讀

DAEDALUS 是一種從自生成練習中建立可重用智能體記憶的方法,無須既有任務或 oracle 驗證器。在 AppWorld、τ^2-bench 及 AutomationBench 上,相較無記憶基線,平均成功率最多提升 15.9 個百分點,pass^5 最多提升 2.2x,推理成本低於大部分使用訓練任務的方法。

正文

Published on Oct 6

Authors:

,

,

,

Abstract

LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, τ^2-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: www.github.com/illuin-tech/daedalus.

View arXiv page View PDF Project page GitHub 0 Add to collection

Get this paper in your agent:

hf papers read 2610.08048

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.08048 in a model README.md to link it from this page.

Datasets citing this paper 1

illuin/daedalus-traces

Viewer •

Updated about 3 hours ago

•

20

•

2

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.08048 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co