跳到正文
原文
Zyphra Research·· 6 小時前AI 評分59

Zyphra Research 發現 GPT 式 Transformer 在持續及平穩學習中出現可塑性損失

ResearchPlasticity Loss in Continual LearningZyphra Research studies the loss of the ability to learn after continued training—loss of plasticity—in GPT-style decoder-only Transformers. We discover that models with 5M to 314M non-embedding parameters lose plasticity when trained on a multilingual continual learning problem, as measured by deterioration on a held-out probing task.We derive a scaling law predicting that the onset of plasticity loss scales sublinearly with the number of non-embedding parameters. This prediction implies that scaling the number of parameters to prevent plasticity loss has diminishing returns and may be an inefficient approach to addressing the problem. In addition, we demonstrate that plasticity loss also occurs during stationary learning, suggesting that current approaches to training large language models are similarly susceptible to plasticity loss. Our work indicates that loss of plasticity might be a key bottleneck for full continual learning of LLMs.

AI 導讀

Zyphra Research 發現,非嵌入參數量介乎 5M 至 314M 的 GPT 式 Transformer,在多語言持續學習中會出現可塑性損失。研究提出縮放定律 T = 1.3 × 10^-5 · P^0.8269,預測可塑性損失的出現時間隨非嵌入參數量呈次線性增長,增加參數量的延後作用會遞減。研究亦在平穩學習中觀察到相似現象,表明可塑性損失並非只見於持續學習設定。

來源:Zyphra Research · zyphra.com