计算机科学
资源配置
推论
资源管理(计算)
人工智能
资源(消歧)
数据建模
分布式计算
数据挖掘
机器学习
钥匙(锁)
自适应系统
调度(生产过程)
计算机网络
作者
Guan Wang,Wenhao Fan,Xiao Li,Guangtao Zhou,Liang Xin,Zhaoyang Yu,Yuanan Liu
标识
DOI:10.1109/tmc.2026.3673729
摘要
By reducing the size of transmitted data between device-side and edge-side machine learning model parts, intermediate activation (IA) compression can alleviate communication overhead, lower latency, and conserve energy, thus enhancing the model partitioning in edge-device collaborative inference scenarios. However, existing studies lack refined resource allocation and fail to jointly optimize edge-device communication and computing resources, especially in terms of IA compression rates, leading to sub-optimal inference performance. To this end, we propose an adaptive model partitioning, intermediate activation compression, and resource allocation scheme for efficient edge-device collaborative inference. We jointly optimize the model partitioning point selection, IA compression rates control, computing resource allocation for both edge and devices, and device transmission power allocation. Our goal is to minimize the weighted sum of inference accuracy loss, inference latency, and device energy consumption. To solve the high-complexity optimization problem efficiently, we design a DRL-based algorithm, which decouples the problem into sub-problems firstly, and then employs an SD3 (Softmax Deep Double Deterministic Policy Gradients)-based DRL method to solve the partitioning point and IA compression rates sub-problem, and utilizes various numerical methods to solve the sub-problems of local and edge computing resource allocation and transmission power control. Extensive comparative simulations with four different schemes under different environmental parameters demonstrate the superiority and robustness of our approach.
科研通智能强力驱动
Strongly Powered by AbleSci AI