跳到正文
HuggingFace Daily Papers(社區熱門論文)·· 2 天前AI 評分40

以純量伴隨匹配進行 Q-learning

Q-Learning with Scalar Adjoint Matching

AI 導讀

研究提出 Q-learning with Scalar Adjoint Matching(SQAM),結合純量伴隨匹配與政策生成動作上的價值懲罰,以微調 flow policy。在 OGBench 最具挑戰的四個領域,SQAM 的成功率較各領域最強基線高 18 至 35 個百分點。在真實雙臂機械人上微調視覺語言動作政策時,SQAM 在全部三項任務的表現均優於監督式微調。

正文

Published on Oct 7

Authors:

,

,

,

,

Abstract

Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian product through the policy at every step, a cost that grows with the number of flow steps and the policy size. We observe that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this finding, we derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time, eliminating the per-step vector--Jacobian products. We further find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. Based on these findings, we propose Q-learning with Scalar Adjoint Matching (SQAM), which combines the scalar adjoint with a value penalty at those actions. SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether SQAM extends to large pretrained policies, we also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.

View arXiv page View PDF Project page GitHub 3 Add to collection

Get this paper in your agent:

hf papers read 2610.10437

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.10437 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.10437 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.10437 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co