跳到正文
HuggingFace Daily Papers(社區熱門論文)·· 13 天前AI 評分51

JumpStart 從 160,000 次訓練中提煉離線策略學習經驗

JumpStart Your Policy Learning with Lessons from 160,000 Training Runs

AI 導讀

這項離線強化學習與模仿學習研究在 114 個數據集上訓練逾 160,000 個策略,發現沒有任何算法在所有環境中佔優。研究指出,超參數調校會頻繁改變算法排名,基準組成也可能導致互相矛盾的結論。研究亦發佈 JumpStart 資源套件,收錄所有已訓練策略、各模型的分數和超參數、各環境的強基線、訓練及評估程式碼,以及供查閲、分析和貢獻結果的可擴展網站,並提供按數據集推薦算法的工具。

正文

Published on Sep 25

Authors:

,

,

,

Abstract

Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has shown that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties, but these sources of variability have not been systematically investigated together at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 160,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance among the strongest methods is often close, but the leaders differ substantially across environments. We find that proper hyperparameter tuning frequently reshuffles perceived algorithm rankings and that benchmark composition can produce conflicting conclusions. We also study hyperparameter sensitivity and transfer across environments, identifying a simple strategy for deriving strong default configurations. We use our findings to develop a dataset-conditioned recommender that provides task-specific algorithm recommendations for practitioners. Finally, we release JumpStart: a resource suite containing every trained policy, per-model scores and hyperparameters, strong baselines across all environments, training and evaluation code, and an extensible website for retrieving, analyzing, and contributing results. Together, these resources aim to make offline policy-learning research more reliable and enable future work beyond the scope of this study.

View arXiv page View PDF Project page GitHub Add to collection

Get this paper in your agent:

hf papers read 2609.13730

Don't have the latest CLI?

curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.13730 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.13730 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.13730 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.

來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co