Goodfire Research·· 6 小時前AI 評分57
Goodfire Research 發佈 Predictive Data Debugging 研究,訓練前預測並調整模型從數據中學到的內容
Predictive Data Debugging: Reveal and Shape What Your Model Learns, Before You Train
AI 導讀
Goodfire Research 發佈 Predictive Data Debugging 研究,提出在訓練前預測偏好數據如何影響模型行為的方法。其預測與模型實際學到的結果達到 R² = 0.9,並可追溯至相關數據羣組。使用 Dolci 和 Tulu 3 數據集的案例顯示,DPO 可能削弱模型對部分有害請求的穩健性;以除錯數據訓練則可同時改善安全與效能。
來源:Goodfire Research · goodfire.com