计算机科学
建筑
计算机体系结构
深度学习
嵌入式系统
炸薯条
内存带宽
芯片上的系统
软件
设计流量
带宽(计算)
计算机硬件
人工智能
操作系统
电信
艺术
视觉艺术
出处
期刊:IEEE Micro
[Institute of Electrical and Electronics Engineers]
日期:2023-05-01
卷期号:43 (3): 18-30
被引量:28
标识
DOI:10.1109/mm.2023.3256384
摘要
The compute and memory demands for deep learning and machine learning (ML) have increased by several orders of magnitude in just the last couple of years, and there is no end in sight. Traditional improvements in processor performance alone struggle to keep up with the exponential demand. A new chip architecture co-designed with the ML algorithms can be better equipped to satisfy this unprecedented demand and enable the ML workloads of the future. This article describes the Cerebras architecture and how it is designed specifically with this purpose, from the ground up, as a wafer-sized chip to enable emerging extreme-scale ML models. It uses fine-grained data flow compute cores to accelerate unstructured sparsity, distributed static random-access memory for full memory bandwidth to the data paths, and a specially designed on-chip and off-chip interconnect for ML training. With these techniques, the Cerebras architecture provides unique capabilities beyond traditional designs.
科研通智能强力驱动
Strongly Powered by AbleSci AI