DRL-Based Adaptive Model Partitioning, Intermediate Activation Compression, and Resource Allocation for Edge-Device Collaborative Inference

计算机科学 资源配置 推论 资源管理(计算) 人工智能 资源(消歧) 数据建模 分布式计算 数据挖掘 机器学习 钥匙(锁) 自适应系统 调度(生产过程) 计算机网络
作者
Guan Wang,Wenhao Fan,Xiao Li,Guangtao Zhou,Liang Xin,Zhaoyang Yu,Yuanan Liu
出处
期刊:IEEE Transactions on Mobile Computing [IEEE Computer Society]
卷期号:25 (8): 12704-12721
标识
DOI:10.1109/tmc.2026.3673729
摘要

By reducing the size of transmitted data between device-side and edge-side machine learning model parts, intermediate activation (IA) compression can alleviate communication overhead, lower latency, and conserve energy, thus enhancing the model partitioning in edge-device collaborative inference scenarios. However, existing studies lack refined resource allocation and fail to jointly optimize edge-device communication and computing resources, especially in terms of IA compression rates, leading to sub-optimal inference performance. To this end, we propose an adaptive model partitioning, intermediate activation compression, and resource allocation scheme for efficient edge-device collaborative inference. We jointly optimize the model partitioning point selection, IA compression rates control, computing resource allocation for both edge and devices, and device transmission power allocation. Our goal is to minimize the weighted sum of inference accuracy loss, inference latency, and device energy consumption. To solve the high-complexity optimization problem efficiently, we design a DRL-based algorithm, which decouples the problem into sub-problems firstly, and then employs an SD3 (Softmax Deep Double Deterministic Policy Gradients)-based DRL method to solve the partitioning point and IA compression rates sub-problem, and utilizes various numerical methods to solve the sub-problems of local and edge computing resource allocation and transmission power control. Extensive comparative simulations with four different schemes under different environmental parameters demonstrate the superiority and robustness of our approach.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
hellohql完成签到,获得积分10
1秒前
阳光的衬衫完成签到,获得积分10
1秒前
卡皮巴拉完成签到 ,获得积分10
2秒前
震动的君浩完成签到,获得积分10
3秒前
3秒前
4秒前
4秒前
胡杨完成签到,获得积分20
5秒前
5秒前
bkagyin的应助被ztt采纳,获得10
6秒前
tangzhidi发布了新的文献求助10
6秒前
7秒前
科研通AI6.2的应助被Keats采纳,获得10
8秒前
林中雀完成签到 ,获得积分10
8秒前
卡卡发布了新的文献求助10
9秒前
HIPPOOOO发布了新的文献求助10
9秒前
9秒前
111发布了新的文献求助10
9秒前
夕青完成签到 ,获得积分10
10秒前
Lei发布了新的文献求助10
10秒前
zhaozhenyang发布了新的文献求助10
10秒前
吴浩南完成签到,获得积分10
10秒前
10秒前
11秒前
SDD发布了新的文献求助10
11秒前
11秒前
long0809发布了新的文献求助10
11秒前
Lucas的应助被直率雪曼采纳,获得10
12秒前
芽芽的应助被无私藏鸟采纳,获得10
12秒前
田様的应助被天天开心采纳,获得10
13秒前
13秒前
熊噗噗完成签到,获得积分10
14秒前
Future发布了新的文献求助10
14秒前
15秒前
jxf发布了新的文献求助10
16秒前
那只幸运的小肥羊完成签到,获得积分10
17秒前
Hello的应助被淡蓝色采纳,获得10
17秒前
18秒前
想吃泡粉发布了新的文献求助30
18秒前
小白发布了新的文献求助10
18秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
Organizational Behavior 510
Arbitrage Theory in Discrete and Continuous Time 500
Fortepian Chopina 400
A Silent Apostrophe:The Fayum Portraits 310
四川大学学位论文.郭瑞昂. 基于高压热扩散的n型磷掺杂金刚石半导体制备研究 300
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 有机化学 化学工程 内科学 物理 生物化学 复合材料 催化作用 细胞生物学 人工智能 心理学 无机化学 基因 遗传学
热门帖子
关注 科研通微信公众号,转发送积分 7832151
求助须知:如何正确求助?哪些是违规求助? 9356044
关于积分的说明 20586481
捐赠科研通 7424502
什么是DOI,文献DOI怎么找? 3336822
关于科研通互助平台的介绍 2481346
邀请新用户注册赠送积分活动 2357505