亲爱的研友该休息了!由于当前在线用户较少,发布求助请尽量完整地填写文献信息,科研通机器人24小时在线,伴您度过漫漫科研夜!身体可是革命的本钱,早点休息,好梦!

MaPHeA: A Framework for Lightweight Memory Hierarchy-aware Profile-guided Heap Allocation

计算机科学 堆(数据结构) 内存层次结构 德拉姆 仿形(计算机编程) 操作系统 隐藏物 地点 嵌入式系统 覆盖 并行计算 分布式计算 计算机硬件 程序设计语言 语言学 哲学
作者
Deok-Jae Oh,Yaebin Moon,Do Kyu Ham,Tae Jun Ham,Yongjun Park,Jae W. Lee,Jung Ho Ahn,Eojin Lee
出处
期刊:ACM Transactions in Embedded Computing Systems [Association for Computing Machinery]
卷期号:22 (1): 1-28 被引量:6
标识
DOI:10.1145/3527853
摘要

Hardware performance monitoring units (PMUs) are a standard feature in modern microprocessors, providing a rich set of microarchitectural event samplers. Recently, numerous profile-guided optimization (PGO) frameworks have exploited them to feature much lower profiling overhead compared to conventional instrumentation-based frameworks. However, existing PGO frameworks mainly focus on optimizing the layout of binaries; they overlook rich information provided by the PMU about data access behaviors over the memory hierarchy. Thus, we propose MaPHeA, a lightweight M emory hierarchy- a ware P rofile-guided He ap A llocation framework applicable to both HPC and embedded systems. MaPHeA guides and applies the optimized allocation of dynamically allocated heap objects with very low profiling overhead and without additional user intervention to improve application performance. To demonstrate the effectiveness of MaPHeA, we apply it to optimizing heap object allocation in an emerging DRAM-NVM heterogeneous memory system (HMS), selective huge-page utilization, and controlling the cacheability of the objects with the low temporal locality. In an HMS, by identifying and placing frequently accessed heap objects to the fast DRAM region, MaPHeA improves the performance of memory-intensive graph-processing and Redis workloads by 56.0% on average over the default configuration that uses DRAM as a hardware-managed cache of slow NVM. By identifying large heap objects that cause frequent TLB misses and allocating them to huge pages, MaPHeA increases the performance of the read and update operations of Redis by 10.6% over the transparent huge-page implementation of Linux. Also, by distinguishing the objects that cause cache pollution due to their low temporal locality and applying write-combining to them, MaPHeA improves the performance of STREAM and RADIX workloads by 20.0% on average over the system without cacheability control.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
4秒前
cc完成签到,获得积分10
10秒前
FashionBoy应助lzx采纳,获得10
17秒前
17秒前
科研通AI6.4应助俏皮幻悲采纳,获得10
27秒前
29秒前
38秒前
42秒前
56秒前
56秒前
1分钟前
小二郎应助科研通管家采纳,获得10
1分钟前
1分钟前
1分钟前
1分钟前
俏皮幻悲发布了新的文献求助10
1分钟前
lzx发布了新的文献求助10
1分钟前
1分钟前
1分钟前
JamesPei应助Giselle采纳,获得10
1分钟前
2分钟前
2分钟前
砍瓜切菜发布了新的文献求助10
2分钟前
2分钟前
2分钟前
Giselle发布了新的文献求助10
2分钟前
2分钟前
Ava应助自由书文采纳,获得10
2分钟前
Giselle完成签到,获得积分10
2分钟前
2分钟前
橙汁完成签到,获得积分10
3分钟前
自由书文发布了新的文献求助10
3分钟前
flysteven92完成签到 ,获得积分10
3分钟前
3分钟前
3分钟前
搜集达人应助科研通管家采纳,获得10
3分钟前
俏皮幻悲发布了新的文献求助10
3分钟前
3分钟前
LV完成签到,获得积分20
3分钟前
3分钟前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
2026年中国辛酸癸酸聚乙二醇甘油酯行业市场现状调查及投资机会研判报告 1000
模型平均及其应用 900
Nondestructive Testing Handbook: Vol. 4, Thermal and Infrared Testing (IR), 4th ed 800
Évora na Idade Média 555
作者名:Kristopher P. Plain,悉尼大学的,目前只能查到其四篇论文,想找到其博士论文 550
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7346567
求助须知:如何正确求助?哪些是违规求助? 8958697
关于积分的说明 19023783
捐赠科研通 6997345
什么是DOI,文献DOI怎么找? 3220101
关于科研通互助平台的介绍 2385047
邀请新用户注册赠送积分活动 2200360