Zyphra Research·· 5 小時前AI 評分43
Zyphra 提出 Online Vector-Quantized (OVQ) Attention 處理長上下文
ResearchOnline Vector Quantized AttentionIn this blog, we describe a novel sequence mixing layer developed here at Zyphra that aims to find a better compromise between memory-compute costs and long-context capabilities than standard sequence mixing layers. We call this layer Online Vector-Quantized (OVQ) attention.February 5, 2026
AI 導讀
Zyphra 提出 Online Vector-Quantized(OVQ)attention,以線性計算複雜度和常數記憶體複雜度處理長上下文。其稀疏狀態更新可擴大記憶容量;實驗顯示,在 16k+ 上下文長度下,OVQ-attention 使用的記憶狀態僅為強自注意力基線的一小部分,表現相若或略有差距。長上下文任務中,OVQ-attention 顯著勝過線性注意力及 SSM 基線。
來源:Zyphra Research · zyphra.com