EngramEdit:透過條件記憶解耦大語言模型的知識更新
EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
EngramEdit 透過條件記憶解耦大語言模型的事實知識更新,並在不改動 Transformer 主幹的情況下進行編輯。實驗顯示,其知識編輯成功率接近完美;更新後的知識可用於未見過的表述及多跳推理,在鏈式推理(CoT)提示下,準確率接近最強基線的 3 倍。隨事實更新累積,無關知識和一般能力大致得以保留。
Published on Oct 7
Authors:
,
,
,
,
,
Abstract
Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.
View arXiv page View PDF Project page GitHub Add to collection
來源:HuggingFace Daily Papers(社區熱門論文) · huggingface.co