跳到正文
原文
Goodfire Research·· 6 小時前AI 評分61

Goodfire Research 發現模型內部訊號可用於大規模偵測 reward hacking

Models know when they’re reward hacking — and we can catch them at scale

AI 導讀

Goodfire Research 發現,開源模型內部存在與 reward hacking 相關的訊號,activation probes 可用於大規模偵測這類行為。

來源:Goodfire Research · goodfire.com