延迟(音频)
计算机科学
高效能源利用
并行计算
电阻随机存取存储器
浮点型
乘法(音乐)
计算机硬件
操作系统
电气工程
工程类
物理
电信
电压
声学
作者
Xianwu Hu,Yu Wang,Zizhao Ma,Gan Wen,Zeming Wang,Zhichao Lu,Yunlong Liu,Yanlei Li,Xingdong Liang,Xiaoyang Zeng,Yufeng Xie
出处
期刊:IEEE Transactions on Circuits and Systems Ii-express Briefs
[Institute of Electrical and Electronics Engineers]
日期:2023-06-07
卷期号:70 (11): 4216-4220
被引量:8
标识
DOI:10.1109/tcsii.2023.3283418
摘要
High-precision computation with low latency and high energy efficiency is required for AI-driven application and scientific computing. Emerging compute-in-memory (CIM) technology shows a great potential to accelerate multiplication and accumulation (MAC) operations which are frequently executed in such scenarios. Resistive RAM (RRAM) is highly suitable for CIM due to its excellent features such as nonvolatility, small cell size and MAC-friendly structure. However, the existing RRAM CIMs focus on the acceleration of fixed-point/integer operations. Several works adopt the logic-CIM structure to support high-precision Floating-point (FP) calculations, but they require lots of cycles and area to perform a FP operation. To meet the need of low latency and high energy efficiency of widely used FP calculation, we propose an accelerated FP-MAC architecture, based on 40nm RRAM CIM array. A full-parallel data input scheme and triangle weights arrangement is proposed for low latency multi-bits multiplication. A non-uniformly grouped sense amplifiers (NUGSAs) array is adopted for energy and area saving. Experiments show that the proposed FP-MAC design achieves an energy efficiency of up to 8.8 TFLOPS/W at FP8 mode and 3.3 TFLOPS/W at bFP16 mode, and the computing latency is 3.34ns.
科研通智能强力驱动
Strongly Powered by AbleSci AI