环境科学
拦截
遥感
计算机科学
气象学
噪音(视频)
地质学
雷达跟踪器
期限(时间)
鉴定(生物学)
作者
He Cai,Yibo Zhang,Youfeng Su,Huanli Gao
标识
DOI:10.1109/tnnls.2026.3706812
摘要
This article studies a swarm-to-swarm interception problem, where a swarm of intercept uncrewed aerial vehicles (UAVs) attempt to intercept a swarm of target UAVs based on pure vision-feedback. The proposed interception strategy employs a hierarchical structure consisting of three parts, namely, image processing, target allocation, and motion planning. As the core of the interception strategy, the target allocation model is trained by a novel dynamic sampling multiagent deep deterministic policy gradient (DS-MADDPG) algorithm, which, different from other classic MADDPG algorithms, features three specialties: the introduction of independent replay buffers and dynamic sampling networks to optimize learning efficiency; the incorporation of multidimensional convolution structures and a multihead attention mechanism into the policy network to capture the spatial and temporal relationships between agents and extract features; the design of a TD3-inspired twin critic architecture to mitigate overestimation bias in action-value functions. The performance of the proposed interception strategy is verified by comprehensive case studies in comparison with other existing algorithms, and the results validate the efficiency and effectiveness of the proposed strategy in terms of multiple performance indices.
科研通智能强力驱动
Strongly Powered by AbleSci AI