跳到正文
原文
HuggingFace Daily Papers(社區熱門論文)·· 5 天前AI 評分41

Magic-W0:面向物理智能的結構化世界—動作基礎模型

Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence

AI 導讀

Magic-W0 是面向物理智能的世界動作基礎模型,聯合建模結構化物理狀態演變與連續動作,並以當前狀態、轉換及未來狀態表示互動。在 RoboDojo-Sim 上,Magic-W0 平均得分為 27.10,在受比較的世界動作模型中最高。在多項真實機械人任務中,模型以有限下游數據微調後仍展現良好表現,支持其泛化及快速適應能力。

正文

Published on Oct 3

Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions. Magic-W0 represents interaction as a Structured World Transition consisting of Current State, Transition, and Future State. Current State combines vision-language context with Current 3D Geometry; Transition is represented by 3D Motion capturing action-induced three-dimensional changes; and Future State is represented by Future Semantics describing task-relevant outcomes. To couple prediction and control, we propose a layer-aligned world-action interaction architecture in which evolving action hypotheses condition world-transition prediction, while predicted world representations continuously inform action generation. Magic-W0 is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision for geometry, 3D motion, and future semantics from pre-trained visual models. Inference-time interventions show that structured world representations respond systematically to changes in candidate actions and that action-related information propagates through shared 3D representations into future semantic predictions. On RoboDojo-Sim, Magic-W0 achieves an average Score of 27.10, the highest among the compared WAMs. Across multiple real-robot tasks, it also demonstrates strong downstream performance after fine-tuning with limited downstream data, supporting generalization and rapid adaptation.

View arXiv page View PDF Project page GitHub 2 Add to collection

Get this paper in your agent:

hf papers read 2609.39870

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.39870 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.39870 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.39870 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co