电阻随机存取存储器
静态随机存取存储器
延迟(音频)
GSM演进的增强数据速率
光电子学
材料科学
计算机科学
电气工程
计算机硬件
工程类
电信
电压
作者
Chen Mu,Z.-X. Huang,H. W. Jiang,Jie Liao,Yuliang Zhou,Liang Chen,Yechu Zhang,Haozhe Zhu,Jianguo Yang,Qi Liu,Chixiao Chen
出处
期刊:IEEE Journal of Solid-state Circuits
[Institute of Electrical and Electronics Engineers]
日期:2025-06-17
卷期号:60 (10): 3626-3638
被引量:4
标识
DOI:10.1109/jssc.2025.3577335
摘要
The resistive random access memory (RRAM)-based computing-in-memory features high density and high energy efficiency on edge. However, fine-tuning RRAM-based SoCs remains challenging due to the inherent limitations of non-volatile memory (NVM) characteristics. This work proposes an NVM-endurance/latency-aware collaborative RRAM/static random access memory (SRAM) compute-in-memory (CIM) accelerator that addresses the difficulties associated with RRAM write operations during fine-tuning tasks for both CNN and transformer models. The major contributions are: 1) RRAM-most significant bit (MSB)–SRAM-least significant bit (LSB) (RMSL)-based collaborative CIM macros, mitigating RRAM cell flipping times and alleviating the endurance concern; 2) an RRAM-sparse-SRAM-dense (RSSD) weight updating engine, minimizing the long reading and writing latency associated with RRAM access; and 3) a row-wise pipeline weight gradient (WG) computing data flow with low-hardware overhead. With a maximum bit update of ten times per fine-tuning (20 epochs, from start to convergence), the system achieves an energy efficiency of 76.25 TOPS/W for the CIM macro and 22.07 TOPS/W for the overall system. Thanks to the bit-level collaborative CIM, the RRAM CIM macro achieves the same hardware utilization during fine-tuning and inference processes to support RRAM’s energy-efficient computing. The proposed CIM accelerator, fabricated using 28-nm CMOS technology, achieves up to $143{\times }$ RRAM endurance improvement, $117{\times }$ and $144{\times }$ reduction in RRAM write power and latency.
科研通智能强力驱动
Strongly Powered by AbleSci AI