干扰
强化学习
雷达
计算机科学
跳频扩频
马尔可夫决策过程
雷达干扰与欺骗
背景(考古学)
马尔可夫过程
人工智能
脉冲多普勒雷达
电信
数学
雷达成像
统计
物理
热力学
古生物学
生物
标识
DOI:10.1109/radarconf2043947.2020.9266402
摘要
It is shown that frequency hopping and pulsewidth allocation strategy can provide enhanced anti-jamming performance for the radar systems. The current anti-jamming methods often have difficulty in adapting their policy to the complicated and unpredictable jamming environment. To address this limitation, a reinforcement learning-based joint adaptive frequency hopping and pulse-width allocation scheme is proposed. By applying the reinforcement learning, the radar can learn the optimized anti-jamming policy by interacting with the environment and requires little prior information. In the proposed scheme, we first establish a reward model to quantify the performance of radar anti-jamming decisions. Then, the radar anti-jamming decision process is modeled as a Markov decision process. As one of the widely-used reinforcement learning algorithms, the Q-learning, which can converge to the optimized policy with probability 1, is utilized to learn the optimized radar anti-jamming policy in the context of lacking a perfect environmental knowledge. Numerical results are shown to verify the effectiveness of our proposed strategy.
科研通智能强力驱动
Strongly Powered by AbleSci AI