iADD:改善擴散策略最佳化中的對齊與多樣性
iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
iADD 提出以增量 Feynman-Kac 訓練為基礎的擴散模型策略最佳化方法,以改善對齊與多樣性的取捨。理論分析指出,僅更新較後期時間步可能損害多樣性,與先前研究的結論相反。研究在三項不同任務中將方法與相關方法比較,並以各組件消融實驗驗證其在對齊和多樣性方面均有性能提升。
Published on Oct 1
Authors:
,
,
,
Abstract
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that only-latter timestep updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
View arXiv page View PDF Project page GitHub 0 Add to collection
Get this paper in your agent:
hf papers read 2610.01789
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2610.01789 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2610.01789 in a dataset README.md to link it from this page.
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co