跳到正文
原文
Prime Intellect·· 6 小時前AI 評分63

Prime Intellect 研究 reward hacking 的梯度動態,並推出 Sprints 社羣研究計劃

ResearchMAY 20TH, 2026Systematic Reward Hacking and Prime Sprints

AI 導讀

Prime Intellect 以 Llama 3.2-1B-Instruct 進行受控 RL 實驗,指出 reward hacking 是梯度動態問題,受可見獎勵與隱藏獎勵的競爭影響。

來源:Prime Intellect · primeintellect.ai