计算机科学
加速
背景(考古学)
推论
散列函数
时间复杂性
算法
基质(化学分析)
秩(图论)
矩阵完成
理论计算机科学
二次方程
并行计算
数学
人工智能
复合材料
几何学
古生物学
量子力学
组合数学
高斯分布
生物
物理
计算机安全
材料科学
作者
In‐Su Han,Rajesh Jayaram,Amin Karbasi,Vahab Mirrokni,David P. Woodruff,Amir Zandieh
标识
DOI:10.48550/arxiv.2310.05869
摘要
We present an approximate attention mechanism named HyperAttention to address the computational challenges posed by the growing complexity of long contexts used in Large Language Models (LLMs). Recent work suggests that in the worst-case scenario, quadratic time is necessary unless the entries of the attention matrix are bounded or the matrix has low stable rank. We introduce two parameters which measure: (1) the max column norm in the normalized attention matrix, and (2) the ratio of row norms in the unnormalized attention matrix after detecting and removing large entries. We use these fine-grained parameters to capture the hardness of the problem. Despite previous lower bounds, we are able to achieve a linear time sampling algorithm even when the matrix has unbounded entries or a large stable rank, provided the above parameters are small. HyperAttention features a modular design that easily accommodates integration of other fast low-level implementations, particularly FlashAttention. Empirically, employing Locality Sensitive Hashing (LSH) to identify large entries, HyperAttention outperforms existing methods, giving significant speed improvements compared to state-of-the-art solutions like FlashAttention. We validate the empirical performance of HyperAttention on a variety of different long-context length datasets. For example, HyperAttention makes the inference time of ChatGLM2 50\% faster on 32k context length while perplexity increases from 5.6 to 6.3. On larger context length, e.g., 131k, with causal masking, HyperAttention offers 5-fold speedup on a single attention layer.
科研通智能强力驱动
Strongly Powered by AbleSci AI