跳到正文
原文
HuggingFace Daily Papers(社區熱門論文)·· 4 天前AI 評分35

DiVeR:面向 VLA 測試時擴展的決策關鍵性驗證器學習

DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling

AI 導讀

DiVeR 提出面向 VLA 測試時擴展的驗證器學習方法,按動作選擇對任務結果的重要程度調整學習權重。方法根據取樣動作表示的分散程度估計決策關鍵性,無需逐步標註或額外環境互動。在 LIBERO、RoboCasa 及 Franka Research 3 機械人的實驗中,DiVeR 均提升任務成功率,驗證器推理額外開銷可忽略。

正文

Published on Oct 4

Authors:

,

,

,

,

,

,

Abstract

Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally, even though their value for candidate discrimination can vary across a trajectory. At many states, plausible actions are similar and provide limited discrimination signal, while only a sparse set of decision-critical states admits meaningfully different actions that can substantially affect downstream outcomes. To address this, we propose DiVeR, which estimates decision criticality from the dispersion of sampled action representations. DiVeR then uses this signal to reweight verifier learning toward states where action selection is most consequential, without requiring step-level annotations or additional environment interaction. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves task success through more effective verifier-guided action selection, while adding negligible verifier inference overhead.

View arXiv page View PDF Add to collection

Get this paper in your agent:

hf papers read 2610.04933

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.04933 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.04933 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.04933 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co