跳到正文
原文
Goodfire Research·· 6 小時前AI 評分57

Goodfire Research 發佈 Predictive Data Debugging 研究,訓練前預測並調整模型從數據中學到的內容

Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train

AI 導讀

Goodfire Research 發佈 Predictive Data Debugging 研究,提出在訓練前預測偏好數據如何影響模型行為的方法。其預測與模型實際學到的結果達到 R² = 0.9,並可追溯至相關數據羣組。使用 Dolci 和 Tulu 3 數據集的案例顯示,DPO 可能削弱模型對部分有害請求的穩健性;以除錯數據訓練則可同時改善安全與效能。

來源:Goodfire Research · goodfire.com