MCaM : Efficient LLM Inference with Multi-tier KV Cache Management

计算机科学 隐藏物 智能缓存 缓存算法 并行计算 缓存失效 CPU缓存 缓存污染 缓存着色 操作系统 计算机网络 页面缓存 异步通信 嵌入式系统 延迟(音频) 重新使用 调度(生产过程) 公共汽车嗅探 吞吐量 MESI协议 德拉姆 管道(软件) 分布式计算
作者
Kexin Chu,Zixu Shen,Sheng-Ru Cheng,Dawei Xiang,Ziqin Liu,Wei Zhang
标识
DOI:10.1109/icdcs63083.2025.00062
摘要

The KV cache in current LLM serving system is primarily used to accelerate processing within a single request and is aggressively deleted once the response is generated. However, in scenarios like virtual assistants and multi-turn conversations, the KV cache can be reused across requests, which can dramatically reduce computation costs and improve serving latency. Caching historical tokens, however, significantly increases memory requirements. Furthermore, existing serving systems treat the request scheduler and KV cache separately, despite their tight coupling.MCaM is a multi-tier cache system that enables the KV cache reuse and sharing across requests. It leverages DRAM as slow- tier memory for storing the KV cache of historical prompts. To efficiently utilize fast-tier Hign Bandwidth Memory(HBM) on GPU, we co-designed the KV cache manager and scheduler to coordinate request scheduling and token placement across tiers. To hide the reload time, MCaM employs a pipeline prefetcher that overlaps communication and computation. Additionally, MCaM incorporates a quality-aware sparsification algorithm to heterogeneously compress the KV cache in each layer. This approach not only reduces data transfer size but also decreases the overall KV cache size. To remove data offloading from a request’s critical path, we designed an asynchronous offload engine that swaps data from HBM to DRAM in the background. Our experiments show that MCaM can reduce TTFT by up to 69% and improve prompt prefilling throughput by 3.3X. It can also reduce the end-to-end latency of LLM inference by up to 58% when request length increase to 4096 tokens.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
DEVIL完成签到,获得积分10
刚刚
aria发布了新的文献求助10
1秒前
2秒前
万能图书馆应助王檬采纳,获得10
2秒前
moya完成签到,获得积分10
2秒前
奋青完成签到 ,获得积分10
2秒前
2秒前
爱读文献的锅包肉完成签到,获得积分10
2秒前
3秒前
3秒前
董星辰发布了新的文献求助10
3秒前
0.0.123完成签到,获得积分10
3秒前
整齐的磬gsq完成签到,获得积分10
3秒前
yjf完成签到 ,获得积分10
3秒前
ghtsmile发布了新的文献求助10
4秒前
xingmeng完成签到,获得积分10
4秒前
4秒前
Yanchen完成签到,获得积分10
4秒前
4秒前
清秋1001完成签到,获得积分10
5秒前
dgqlcc发布了新的文献求助10
5秒前
张红梨完成签到,获得积分10
5秒前
泉竹晓筱完成签到,获得积分10
5秒前
zhaozhao完成签到 ,获得积分10
5秒前
5秒前
科研狗完成签到,获得积分10
5秒前
5秒前
6秒前
7秒前
夜捕白日梦完成签到,获得积分10
7秒前
科研通AI6.4应助吴灵采纳,获得10
7秒前
egg2完成签到,获得积分10
8秒前
8秒前
彭于晏应助hotpig460采纳,获得100
8秒前
面包糠发布了新的文献求助10
9秒前
kstreet0728发布了新的文献求助10
9秒前
铁柱发布了新的文献求助10
9秒前
Leon Lai完成签到,获得积分10
10秒前
仪圆完成签到,获得积分10
10秒前
ZY发布了新的文献求助10
10秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
日本現代怪異事典 副読本 700
悉尼大学博士学位论文,题目:Modelling and testing of one-sided stitched laminated composites. 作者:Kristopher P. Plain 650
Machine Learning for Asset Management and Pricing 600
Numerical analysis of the coupled atmosphere-ocean models (CAO II). II 600
Models for the coupled atmosphere and ocean 600
Évora na Idade Média 555
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7385305
求助须知:如何正确求助?哪些是违规求助? 8992012
关于积分的说明 19128674
捐赠科研通 7022740
什么是DOI,文献DOI怎么找? 3227478
关于科研通互助平台的介绍 2390471
邀请新用户注册赠送积分活动 2208633