Magic-W0:面向物理智能的結構化世界—動作基礎模型
Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence
Magic-W0 是面向物理智能的世界動作基礎模型,聯合建模結構化物理狀態演變與連續動作,並以當前狀態、轉換及未來狀態表示互動。在 RoboDojo-Sim 上,Magic-W0 平均得分為 27.10,在受比較的世界動作模型中最高。在多項真實機械人任務中,模型以有限下游數據微調後仍展現良好表現,支持其泛化及快速適應能力。
Published on Oct 3
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
World-action models (WAMs) augment robot policies with action-conditioned environment dynamics, yet existing approaches largely rely on future observation reconstruction or generic latent prediction and lack structured, control-oriented world representations tightly coupled with action generation. We introduce Magic-W0, a world-action foundation model that jointly models structured physical state evolution and continuous actions. Magic-W0 represents interaction as a Structured World Transition consisting of Current State, Transition, and Future State. Current State combines vision-language context with Current 3D Geometry; Transition is represented by 3D Motion capturing action-induced three-dimensional changes; and Future State is represented by Future Semantics describing task-relevant outcomes. To couple prediction and control, we propose a layer-aligned world-action interaction architecture in which evolving action hypotheses condition world-transition prediction, while predicted world representations continuously inform action generation. Magic-W0 is pre-trained on large-scale egocentric human manipulation, UMI, real-robot, and simulation data, with latent supervision for geometry, 3D motion, and future semantics from pre-trained visual models. Inference-time interventions show that structured world representations respond systematically to changes in candidate actions and that action-related information propagates through shared 3D representations into future semantic predictions. On RoboDojo-Sim, Magic-W0 achieves an average Score of 27.10, the highest among the compared WAMs. Across multiple real-robot tasks, it also demonstrates strong downstream performance after fine-tuning with limited downstream data, supporting generalization and rapid adaptation.
View arXiv page View PDF Project page GitHub 2 Add to collection
Get this paper in your agent:
hf papers read 2609.39870
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.39870 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.39870 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.39870 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co