Lyapunov-Guided Deep Reinforcement Learning for Stable Online Computation Offloading in Mobile-Edge Computing Networks

Lyapunov优化计算机科学移动边缘计算计算卸载帧（网络）在线算法强化学习数学优化边缘计算最优化问题随机优化李雅普诺夫函数无线网络随机规划 GSM演进的增强数据速率无线算法人工智能计算机网络李雅普诺夫方程李雅普诺夫指数数学物理电信非线性系统量子力学混乱的

作者

Suzhi Bi,Liang Huang,Hui Wang,Ying–Jun Angela Zhang

出处

期刊：IEEE Transactions on Wireless Communications [Institute of Electrical and Electronics Engineers]
日期：2021-06-09 卷期号：20 (11): 7519-7537 被引量：209

链接

arxiv.org arxiv.orgdoi.org

标识

DOI：10.1109/twc.2021.3085319

摘要

Opportunistic computation offloading is an effective method to improve the computation performance of mobile-edge computing (MEC) networks under dynamic edge environment. In this paper, we consider a multi-user MEC network with time-varying wireless channels and stochastic user task data arrivals in sequential time frames. In particular, we aim to design an online computation offloading algorithm to maximize the network data processing capability subject to the long-term data queue stability and average power constraints. The online algorithm is practical in the sense that the decisions for each time frame are made without the assumption of knowing the future realizations of random channel conditions and data arrivals. We formulate the problem as a multi-stage stochastic mixed integer non-linear programming (MINLP) problem that jointly determines the binary offloading (each user computes the task either locally or at the edge server) and system resource allocation decisions in sequential time frames. To address the coupling in the decisions of different time frames, we propose a novel framework, named LyDROO, that combines the advantages of Lyapunov optimization and deep reinforcement learning (DRL). Specifically, LyDROO first applies Lyapunov optimization to decouple the multi-stage stochastic MINLP into deterministic per-frame MINLP subproblems. By doing so, it guarantees to satisfy all the long-term constraints by solving the per-frame subproblems that are much smaller in size. Then, LyDROO integrates model-based optimization and model-free DRL to solve the per-frame MINLP problems with very low computational complexity. Simulation results show that under various network setups, the proposed LyDROO achieves optimal computation performance while stabilizing all queues in the system. Besides, it induces very low computation time that is particularly suitable for real-time implementation in fast fading environments.

求助该文献

最长约 10秒，即可获得该文献文件

Lyapunov-Guided Deep Reinforcement Learning for Stable Online Computation Offloading in Mobile-Edge Computing Networks

今日热心研友