跳到正文
原文
HuggingFace Daily Papers(社區熱門論文)·· 1 天前AI 評分42

Attacca:狀態連續條件下長時程具身智能體的目標導向控制

Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents

AI 導讀

Attacca 提出以完整搜尋至互動軌跡訓練視覺目標條件式策略的方法,應對具身智能體長時程任務中的狀態連續問題。該方法在 Minecraft 多項短、長時程任務中取得 39.0-47.5% clean success,較最強基線提高 1.7-2.4x。長時程任務完成率為 54%、30% 和 28%,最高提升 7x。

正文

Published on Oct 6

Authors:

Abstract

A central capability of embodied agents is to accomplish complex objectives through sequences of interdependent tasks. Yet existing visual goal-conditioned policies underlying these agents are typically evaluated on isolated interactions where the target is already visible, and thus do not capture the conditions that arise during continuous long-horizon task execution. In such settings, each task begins from the state left by the previous one: the agent may end at a different position and orientation, the world may have been modified, and the next interaction target may lie outside the current field of view. As a result, agents relying on such policies may struggle to proceed to the next task when they cannot ground their target in the current observation. To address this challenge, we propose Attacca, a new approach that trains visual goal-conditioned policies on complete search-to-interact trajectories using goal images decoupled from the execution environment. Attacca uses context-decoupled goal sampling to pair each demonstration with a class-compatible masked goal image from another world, removing direct scene and pose correspondence. It learns dense current-view grounding through a target-mask prediction head, providing auxiliary supervision beyond action imitation. We further introduce behavioral-phase conditioning that teaches the policy to distinguish Search, Approach, and Interact stages and adapt its control as execution progresses. We evaluate Attacca on multiple short- and long-horizon embodied tasks in Minecraft. Our method achieves 39.0-47.5% clean success, improving over the strongest baseline by 1.7-2.4x. On long-horizon tasks, it attains 54%, 30%, and 28% completion, yielding up to a 7x improvement.

View arXiv page View PDF Project page GitHub Add to collection

Get this paper in your agent:

hf papers read 2610.07785

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

willsuh/attacca

0.2B •

Updated about 2 hours ago

•

12

Datasets citing this paper 1

willsuh/attacca-dataset

Updated about 2 hours ago

•

35

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.07785 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co