Data Convection

作者
Soheil Khadirsharbiyani,Jagadish Kotra,Karthik Rao,Mahmut Kandemir
出处
期刊:Proceedings of the ACM on measurement and analysis of computing systems [Association for Computing Machinery]
卷期号:6 (1): 1-25 被引量:3
标识
DOI:10.1145/3508027
摘要

Stacked DRAMs have been studied, evaluated in multiple scenarios, and even productized in the last decade. The large available bandwidth they offer make them an attractive choice, particularly, in high-performance computing (HPC) environments. Consequently, many prior research efforts have studied and evaluated 3D stacked DRAM-based designs. Despite offering high bandwidth, stacked DRAMs are severely constrained by the overall memory capacity offered. In this paper, we study and evaluate integrating stacked DRAM on top of a GPU in a 3D manner which in tandem with the 2.5D stacked DRAM increases the capacity and the bandwidth without increasing the package size. This integration of 3D stacked DRAMs aids in satisfying the capacity requirements of emerging workloads like deep learning. Though this vertical 3D integration of stacked DRAMs also increases the total available bandwidth, we observe that the bandwidth offered by these 3D stacked DRAMs is severely limited by the heat generated on the GPU. Based on our experiments on a cycle-level simulator, we make a key observation that the sections of the 3D stacked DRAM that are closer to the GPU have lower retention-times compared to the farther layers of stacked DRAM. This thermal-induced variable retention-times causes certain sections of 3D stacked DRAM to be refreshed more frequently compared to the others, thereby resulting in thermal-induced NUMA paradigms. To alleviate such thermal-induced NUMA behavior, we propose and experimentally evaluate three different incarnations of Data Convection, i.e., Intra-layer, Inter-layer, and Intra + Inter-layer, that aim at placing the most-frequently accessed data in a thermal-induced retention-aware fashion, taking into account both bank-level and channel-level parallelism. Our evaluations on a cycle-level GPU simulator indicate that, in a multi-application scenario, our Intra-layer, Inter-layer and Intra + Inter-layer algorithms improve the overall performance by 1.8%, 11.7%, and 14.4%, respectively, over a baseline that already encompasses 3D+2.5D stacked DRAMs.

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
科目三应助肉丸111采纳,获得10
1秒前
溪风不渡发布了新的文献求助10
1秒前
小殷发布了新的文献求助10
2秒前
3秒前
JamesPei应助简单小鸭子采纳,获得10
4秒前
Judles发布了新的文献求助10
4秒前
Hezhiyong发布了新的文献求助10
5秒前
肥胖的红薯完成签到,获得积分10
5秒前
Orange应助时衍采纳,获得10
5秒前
yuandashazi发布了新的文献求助30
6秒前
6秒前
7秒前
不安的夜柳完成签到 ,获得积分10
8秒前
孟小宝完成签到,获得积分10
8秒前
8秒前
可靠诗蕊关注了科研通微信公众号
8秒前
9秒前
赵赵发布了新的文献求助20
9秒前
Parsee应助cfw采纳,获得10
10秒前
Hezhiyong完成签到,获得积分10
10秒前
10秒前
英俊的铭应助申誉杰采纳,获得10
11秒前
端庄夏天完成签到 ,获得积分10
11秒前
12秒前
Liulu发布了新的文献求助10
13秒前
14秒前
16秒前
小溪发布了新的文献求助10
17秒前
科研通AI6.4应助IchenNG采纳,获得10
17秒前
风中的晓兰完成签到,获得积分10
19秒前
19秒前
王亚茹发布了新的文献求助10
19秒前
20秒前
Hello应助科研通管家采纳,获得10
20秒前
20秒前
aajhajkahna应助科研通管家采纳,获得10
20秒前
20秒前
行稳致远应助科研通管家采纳,获得10
20秒前
aajhajkahna应助科研通管家采纳,获得10
20秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
APA handbook of comparative psychology: Basic concepts, methods, neural substrate, and behavior 1000
Child and Adolescent Mental Health 600
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
Römisch-Germanische Forschungen 500
Electric machines: theory, operating applications, and controls 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7600134
求助须知:如何正确求助?哪些是违规求助? 9176262
关于积分的说明 19648312
捐赠科研通 7176233
什么是DOI,文献DOI怎么找? 3268595
关于科研通互助平台的介绍 2433042
邀请新用户注册赠送积分活动 2262187