跳到正文
熱點事件持續更新

Prime Intellect 發佈獎勵破解研究及實驗環境

1 篇報道1 個報道來源6 小時前更新

先了解這件事

報道摘要

Prime Intellect 以 Llama 3.2-1B-Instruct 進行受控 RL 實驗,指出 reward hacking 是梯度動態問題,受可見獎勵與隱藏獎勵的競爭影響。

摘自 Prime Intellect

報道時間線

沿着報道,瞭解事件的不同側面。

10月6日
  1. Prime Intellect
    Prime Intellect 研究 reward hacking 的梯度動態,並推出 Sprints 社羣研究計劃

    Prime Intellect 以 Llama 3.2-1B-Instruct 進行受控 RL 實驗,指出 reward hacking 是梯度動態問題,受可見獎勵與隱藏獎勵的競爭影響。

本事件熱度走勢

還沒有足夠的連續觀測數據,暫不繪製趨勢。