Lincoln: Real-Time 50~100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash Memory

闪光灯(摄影) 闪存 计算机科学 推论 嵌入式系统 计算机硬件 计算机图形学(图像) 人工智能 艺术 视觉艺术
作者
Weiyi Sun,Mingyu Gao,Zhaoshi Li,Aoyang Zhang,Iris Ying Chou,Jianfeng Zhu,Shaojun Wei,Leibo Liu
标识
DOI:10.1109/hpca61900.2025.00128
摘要

With the widespread use of large language models (LLMs), and with the privacy and cost concerns on cloud-based services, vendors are now pushing LLM inference to consumer devices. However, current attempts only enable real-time inference of low-quality small-sized LLMs. Large-sized LLMs have to load most of their weights from Flash storage for every execution iteration, which dominates the execution time of both the prefill and the generation phase. This performance bottleneck is attributed to both the low internal Flash memory bandwidth and the low transmission bandwidth between Flash and the Neural Processing Unit (NPU). To tackle these two challenges, we present Lincoln, a device-architecture co-design solution with LPDDR-interfaced, Compute-Enabled Flash Memory. On the device level, we boost the Flash internal bandwidth by improving upon existing array shrinking methods, to enable lower read latency and more parallel Flash planes within each Flash die. We specifically leverage 3D hybrid bonding, which is already adopted in consumer Flash products, to maintain high area efficiency and low density loss. On the architecture level, to leverage such increased internal bandwidth for resolving the transmission bottleneck, we propose two solutions for the two distinct phases of LLMs. For the compute-intensive prefill phase, we let Flash devices use the existing high-speed LPDDR interface (originally for DRAM), which offers much higher transmission bandwidth to the NPU than the conventional Flash interface, while maintaining good cost and area efficiency. For the memory-intensive generation phase, we rely on hybrid-bonding-based near-Flash computing to fully utilize the internal Flash bandwidth, and further equip with speculative decoding to eventually reach the real-time latency goal. Our evaluation shows that Lincoln enables real-time inference, with up to $13.23 \times$ and $254.1 \times$ speedups for LLM prefill and generation phases over conventional SSD-based systems.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
顾矜应助胡伟采纳,获得10
1秒前
1秒前
2秒前
kaca完成签到,获得积分10
2秒前
2秒前
3秒前
XY完成签到,获得积分10
3秒前
瓜瓜发布了新的文献求助10
3秒前
kkk发布了新的文献求助10
3秒前
用户123发布了新的文献求助10
3秒前
研友_VZG7GZ应助海贵采纳,获得10
4秒前
搜集达人应助飘逸的天蓝采纳,获得30
4秒前
4秒前
zero完成签到,获得积分10
5秒前
5秒前
6秒前
6秒前
6秒前
kchrisuzad完成签到,获得积分10
7秒前
可爱的函函应助Pami采纳,获得10
7秒前
轻松小伙发布了新的文献求助10
7秒前
麦浪完成签到,获得积分10
8秒前
Jing发布了新的文献求助10
8秒前
赘婿应助Lily采纳,获得10
8秒前
挖机发布了新的文献求助10
9秒前
9秒前
小邋遢完成签到,获得积分10
9秒前
雪白的笑阳发布了新的文献求助100
9秒前
10秒前
小熊猫发布了新的文献求助10
10秒前
11秒前
11秒前
11秒前
11秒前
11秒前
小马甲应助过冷风采纳,获得10
12秒前
12秒前
星星发布了新的文献求助10
12秒前
所所应助欢呼的愚志采纳,获得10
12秒前
千里毅完成签到 ,获得积分10
12秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Navigating Normative Orders. Interdisciplinary Perspectives 800
Organizational Behavior 510
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
CLSI VET01S-2024 Performance Standards for Antimicrobial Disk and Dilution Susceptibility Tests for Bacteria Isolated From Animals (7th Ed) 500
A Case Study on Hotels as Noncongregate Emergency Living Accommodations for Returning Citizens 500
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7757241
求助须知:如何正确求助?哪些是违规求助? 9303638
关于积分的说明 20275298
捐赠科研通 7340757
什么是DOI,文献DOI怎么找? 3311761
关于科研通互助平台的介绍 2462624
邀请新用户注册赠送积分活动 2325463