人工智能
计算机科学
特征(语言学)
RGB颜色模型
特征提取
融合
计算机视觉
传感器融合
模式识别(心理学)
哲学
语言学
作者
Yingxiang Hu,Yanbo Liu,Guo Cao,Jin Wang
标识
DOI:10.1109/tim.2025.3555712
摘要
RGB-T crowd counting methods aim to enhance the counting accuracy of network models under conditions of uneven lighting and low visibility by fusing features from the RGB and thermal modalities. Previous approaches primarily utilized attention mechanisms to extract and fuse complementary RGB and thermal features. However, these methods lack guidance and constraints during the extraction and fusion of multi-modal features and do not fully leverage the complementary advantages between global and local features, leading to suboptimal performance. This paper argues that, by transitioning from global attention to local attention, extracting and fusing the complementary information between global and local multi-modal features can significantly improve the model’s counting performance. To achieve this, we propose an RGB-T crowd counting network based on global-local multimodal feature fusion (GLFNet). Specifically, we first use a multi-head attention mechanism to fuse global multi-modal features and guide the global multi-modal fusion using learnable block-counting guided tokens (BCT). Next, we employ composite spatial attention mechanisms (CSAM) to focus on the local detail information of multi-modal crowd features and facilitate the fusion of local multimodal features. Finally, we utilize a detail contrast loss function (Ld) to capture the complementary advantages between global and local multi-modal features and to guide and constrain the fusion process of multi-modal features. Experimental results on the RGBT-CC and DroneRGBT datasets demonstrate the superior performance of our method.
科研通智能强力驱动
Strongly Powered by AbleSci AI