强化学习
计算机科学
马尔可夫决策过程
能源消耗
调度(生产过程)
增强学习
动态规划
分布式计算
马尔可夫过程
实时计算
数学优化
人工智能
工程类
算法
统计
数学
电气工程
作者
Guanhua Wang,Fang Yang,Jian Song,Zhu Han
标识
DOI:10.1109/tcomm.2023.3347775
摘要
Laser inter-satellite links (LISLs) have greatly extended communication distance between satellites, allowing for establishment of dynamic links to reduce communication delay. However, a closed-loop control is required for LISL, which causes high energy consumption. Proper scheduling of dynamic LISLs can effectively reduce energy consumption and communication delay. In this study, a satellite link mode with three fixed LISLs and one dynamic LISL is designed, and its feasibility is analyzed. The optimization problem is formulated and transformed into a Markov decision process (MDP) by modeling it as a sequential decision. By decomposing states, actions, and reward functions, the MDP is divided into the proposed multi-agent deep reinforcement learning (MADRL). Moreover, compressed sensing is utilized to cut down state information to reduce communication, storage, and computation overhead. Furthermore, network parameters and experience sharing, and prioritized experience replay have been adopted to improve stability and convergence speed of network training with a large number of agents. Experimental results show that under different routing strategies, the proposed MADRL can reduce energy consumption by over 15% and delay by approximately two hops compared to fixed LISLs scenario within several iterations.
科研通智能强力驱动
Strongly Powered by AbleSci AI