跳到正文
原文
HuggingFace Daily Papers(社區熱門論文)·· 20 小時前AI 評分48

AdvSim2Real:在網頁世界模型中訓練網頁智能體抵禦自適應提示詞注入

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

AI 導讀

AdvSim2Real 在凍結的網頁世界模型中共同演化任務課程、注入攻擊者與智能體,訓練網頁智能體抵禦自適應提示詞注入。在 150 項網頁任務中,4B 智能體面對未曾用於訓練的前沿模型攻擊者時,完成率較基礎智能體提升 33.6%。其能力增益亦延伸至真實瀏覽器,且有攻擊和無攻擊時的任務完成率均上升。

正文

Published on Oct 6

Authors:

,

,

,

,

Abstract

Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.

View arXiv page View PDF Project page GitHub Add to collection

Get this paper in your agent:

hf papers read 2610.08773

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.08773 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.08773 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.08773 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co