计算机科学
带宽(计算)
闪光灯(摄影)
与非门
建筑
吞吐量
内存带宽
过程(计算)
随机存取存储器
计算机体系结构
内存体系结构
计算机硬件
嵌入式系统
闪存
闪存模拟器
非易失性存储器
推论
系统体系结构
内存管理
并行计算
高内存
半导体存储器
编码器
计算机工程
功率(物理)
作者
Minho Ha,Euiseok Kim,Hoshik Kim
标识
DOI:10.1109/lca.2026.3660969
摘要
Large language model (LLM) inference requires massive memory capacity to process long sequences, posing a challenge due to the capacity limitations of high bandwidth memory (HBM). High bandwidth flash (HBF) is an emerging memory device based on NAND flash that offers HBM comparable bandwidth with much larger capacity, but suffers from disadvantages such as longer access latency, lower write endurance, and higher power consumption. This paper proposes H3, a hybrid architecture designed to effectively utilize both HBM and HBF by leveraging their respective strengths. By storing read-only data in HBF and other data in HBM, H3 equipped systems can process more requests at once with the same number of GPUs than HBM-only systems, making H3 suitable for gigantic read-only use cases in LLM inference, particularly those employing a shared pre-computed key-value cache. Simulation results show that a GPU system with H3 achieves up to 2.69x higher throughput per power compared to a system with HBM-only. This result validates the cost-effectiveness of H3 for handling LLM inference with gigantic read-only data.
科研通智能强力驱动
Strongly Powered by AbleSci AI