Tree Attention:面向 GPU 叢集長上下文注意力的拓撲感知解碼
ResearchTree Attention: Topology-aware Decoding for Long-Context Attention on GPU clustersZyphra is excited to announce Tree Attention, a novel method for efficiently parallelizing multi-GPU transformer decoding with significant advantages in speed and memory. For instance, we estimate that Tree Attention can decode at the 1M sequence length over 8x faster than existing Ring Attention while requiring 2x less communication volume or more. Moreover, Tree Attention achieves an asymptotic advantage over Ring Attention in the number of devices so the benefit increases dramatically for larger clusters.August 7, 2024
Zyphra 提出 Tree Attention,以樹狀歸約平行化多 GPU Transformer 解碼,實測可提升長上下文注意力的解碼速度並降低記憶體開銷。Zyphra 估算,在 1M 序列長度下,Tree Attention 的解碼速度較 Ring Attention 快逾 8x,通訊量則減至一半或更低。其跨裝置複雜度為對數級,而 Ring Attention 為線性級。
來源:Zyphra Research · zyphra.com