Anthropic:Alignment Science Blog·· 4 天前AI 評分57
這會改變模型答案嗎?CHIVE 以反事實實驗評估真實情境中的大語言模型行為解釋
Would This Change Your Answer? Evaluating Explanations of LLM Behavior in the Wild with Counterfactual Experiments
AI 導讀
Anthropic 提出 CHIVE,以反事實提示詞編輯評估大語言模型行為解釋;三類讀取模型 activation 的工具均未勝過只看逐字稿的基線。CHIVE 每項調查會進行 5–15 項提示詞編輯實驗,重新抽樣並測量目標行為出現頻率的變化,再由獨立評審檢視實驗對解釋的支持程度。
來源:Anthropic:Alignment Science Blog · alignment.anthropic.com