计算机科学
人工神经网络
建筑
人工智能
系统体系结构
特征(语言学)
嵌入式系统
钥匙(锁)
集合(抽象数据类型)
网络体系结构
系统工程
工程类
作者
Feng Jiao,Siying Bi,Shu Wang,Zhong Ma,Jin Huang,Siwei Xiang,Chen Yang
标识
DOI:10.1109/prai67447.2025.11412639
摘要
With the rapid advancement of intelligent technologies and their widespread deployment in edge scenarios, energy efficiency has become a critical challenge for computing platforms. To address the difficulty of improving energy efficiency, this paper proposes a high-efficiency neural network accelerator architecture with dedicated designs for storage and computation, tailored for hybrid-precision workloads. A tightly interleaved storage architecture is developed to enable compact storage and low-latency transmission of multi-precision data, proportionally reducing bandwidth pressure and memory access power consumption under low-bitwidth configurations. Furthermore, a reconfigurable hybrid-precision processing engine (HPPE) is designed, which supports 1/2/4/8-bit dynamic precision through bit-extension and sign-bit compensation techniques, enabling efficient parallel multiplication. The coordinated design of storage and compute subsystems helps mitigate performance bottlenecks. Experimental results on a Xilinx ZCU102 FPGA demonstrate a synthesized frequency of 300 MHz. Compared to a fixed 8-bit design, resource usage increases by only 0.5 % to 23.3 % across different resource types. Running the VGG-16 model on the ImageNet dataset, the proposed design achieves $1.05 \times$ to $7.38 \times$ energy efficiency improvement in 8-bit mode compared to similar accelerators. In 1-bit mode, the compute throughput reaches 10,069.8 GOPS—an $8.21 \times$ super-linear improvement over its own 8-bit mode—and offers $49.12 \times$ to $338.71 \times$ energy efficiency improvement over comparable designs. This work provides a high-efficiency computing foundation for edge-side intelligent systems with dynamic task adaptability.
科研通智能强力驱动
Strongly Powered by AbleSci AI