计算机科学
瓶颈
桥接(联网)
Softmax函数
语言模型
推论
编码(社会科学)
人工智能
解码方法
块(置换群论)
序列(生物学)
变压器
高保真
代表(政治)
加速
循环神经网络
编码(内存)
算法
编码(集合论)
连接主义
机器学习
绩效改进
代码生成
计算机工程
隐藏物
判别式
理论计算机科学
计算复杂性理论
面子(社会学概念)
组分(热力学)
并行计算
作者
Xiuying Wei,Anunay Yadav,Razvan Pascanu,Çaǧlar Gülçehre
标识
DOI:10.48550/arxiv.2507.04416
摘要
Transformers have become the cornerstone of modern large-scale language models, but their reliance on softmax attention poses a computational bottleneck at both training and inference. Recurrent models offer high efficiency, but compressing the full sequence into a fixed-size and holistic representation can suffer from memory degradation in long contexts and limit fine-grained retrieval. To address this, we propose RAT, an intermediate design that bridges the efficiency of RNNs and capacity of attention. RAT partitions the input into chunks, applies recurrence within each chunk for local dependencies, and softmax-based attention across chunks for long-range interactions. This design mitigates memory degradation and enables direct access to distant tokens, while retaining computational efficiency. Empirically, with a chunk size of 16, the RAT block achieves a 7$\times$ improvement in training speed for 100K sequence length and 9$times$ in generation at the 4K position, while maintaining similar performance compared to standard attention. We demonstrate this by training 1.3B parameter models from scratch and performing large-scale evaluations, including short- and long-context benchmarks, as well as supervised fine-tuning~(SFT). We further propose a hybrid architecture that interleaves RAT with local attention. By combining efficient long-range modeling with strong local interactions, this hybrid design not only improves inference speed and reduces cache memory usage, but also consistently enhances performance and shows the overall best results. Code is available at https://github.com/CLAIRE-Labo/RAT.
科研通智能强力驱动
Strongly Powered by AbleSci AI