Goodfire Research·· 5 小時前AI 評分14
時間先驗:語言模型可解釋性中缺失的歸納偏置
Priors in Time: Missing Inductive Biases for Language Model Interpretability
AI 導讀
題為《時間先驗》的文章聚焦語言模型可解釋性中缺失的歸納偏置。頁面另列出以蛋白質嵌入向量改善 AI 智能體生物安全監察、在大規模情況下捕捉模型獎勵破解,以及估算文字生成不確定性動態的研究。頁面亦介紹 Silico,稱其為可解釋性智能體,可用於解釋、除錯及精確控制模型行為。
來源:Goodfire Research · goodfire.com