跳到正文
原文
HuggingFace Daily Papers(社區熱門論文)·· 3 天前AI 評分58

工具使用智能體判定結果無用後,仍很少停止查詢

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

AI 導讀

研究發現,受測七個智能體有 97-100% 的情況把失效來源的結果判為無用,但大多數很少據此停止查詢。只有在實驗框架強制加入整合步驟、要求智能體連續五次判定結果無用後作答時,停止行為才會跟隨證據;這提高了所有模型在失效來源下的成功率,而且預算加倍時停止點仍固定。以 300 條全新問題進行的預先註冊重複實驗,亦確認判斷與停止行為的分離及該規則的效果。

正文

Published on Oct 5

Authors:

,

,

,

,

,

Abstract

An agent whose tool keeps returning nothing useful should stop relying on it. In a retrieval environment with controlled source failures, we separate how agents judge results from what they do. We compare stopping at the same step after longer and shorter runs of results the agent judged useless; this contrast is zero for clock- or deadline-driven stopping. Where we record their judgments, the seven agents we test call a failing source's results useless 97-100% of the time, yet most of them rarely stop on that judgment. Prompt cues change when they stop but not what they stop on. Permission to answer from memory and a reasoning mode can bring early stops regardless of evidence, a stated budget moves the 7-8B models' stops to the deadline, and a stopping rule or call cost in the prompt is followed at most partly. Stopping follows the evidence only when the harness enforces an integration step that makes the agent answer after five consecutive results it judged useless. This step raises failing-source success for every model, keeps the stopping point fixed when the budget doubles, and needs no extra judgment call when the agent states its judgments. A pre-registered replication on 300 fresh questions confirms the dissociation and the rule's effect.

View arXiv page View PDF GitHub 3 Add to collection

Get this paper in your agent:

hf papers read 2610.06191

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.06191 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.06191 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.06191 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co