Lyapunov-Based MADRL Policy in Wireless Powered MEC Assisted Monitoring Systems
计算机科学
无线
电信
作者
Xinying Liu,Yuhan Yi,Wenqian Zhang,Guanglin Zhang
标识
DOI:10.1109/infocomwkshps61880.2024.10620870
摘要
In real-time monitoring systems, wireless de-vices(WDs) are employed for information gathering, with age of information(AoI) serving as a crucial measure to assess the timeliness of status information. The combination of wireless power transfer(WPT) and mobile edge computing(MEC) effectively address the challenges posed by the limited battery lifespan and computation capabilities of WDs. This paper focuses on a WP-MEC system where WDs adopt a zero-waiting strategy. A multi-stage stochastic optimization problem is formulated with the goal of age minimization. By utilizing Lyapunov optimization, we convert the problem into deterministic subproblems. However, solving the subproblems proves challenging as they remain mixed-integer nonlinear programming(MINLP) problems. We propose LyMAPPO and LyIPPO to tackle this problem, which leverage the benefits of Lyapunov optimization and multi-agent deep reinforcement learning(MADRL). MADRL enables WDs to make optimal decisions of task scheduling, resource allocation and computation offloading. Simulation results confirm the findings of our theoretical analyses, demonstrating the improved performance of LyMAPPO and LyIPPO in comparison to bench-mark algorithms.