Goodfire Research·· 5 小時前AI 評分57
Goodfire 剖析 DeepSeek R1 推理模型內部機制,公開兩個 SAE
Under the Hood of a Reasoning Model
AI 導讀
Goodfire 公開兩個以 DeepSeek R1 激活值訓練的稀疏自編碼器(SAE),並分享對推理模型內部機制及操控方式的早期研究。研究發現,R1 的特徵操控需等模型開始輸出「Okay, so the user has asked a question about…」一類思考前綴後才有效;部分特徵被過度操控時,模型反而會回復原有行為。
來源:Goodfire Research · goodfire.com