杠杆(统计)
RGB颜色模型
保险丝(电气)
计算机科学
人工智能
互补性(分子生物学)
解码方法
特征(语言学)
计算机视觉
特征提取
融合
模式识别(心理学)
传感器融合
人工神经网络
融合机制
信息融合
编码(集合论)
数据挖掘
稳健性(进化)
特征向量
机器学习
标识
DOI:10.1109/tim.2025.3618732
摘要
In recent years, crowd counting has received extensive attention due to its wide-ranging applications across various fields. However, the exclusive utilization of red–green–blue (RGB) images for crowd counting is susceptible to adverse factors such as lighting conditions, which significantly impair the accuracy. An increasing number of methods have emerged for crowd counting using both RGB and thermal images. Nevertheless, the underutilization of the complementarity among multi-modal information has hindered the progress of RGB and thermal (RGB-T) crowd counting. In order to fully leverage the complementarity among multi-modal information without consuming excessive computing resources, we propose a Multi-Modal Feature Fusion Network (MMFFNet) based on encoder-fusion-decoder structure for RGB-T crowd counting. We put forward a Feature Fusion Module (FFM) that employs an attention-like approach to effectively fuse the extracted features at multiple levels across RGB and thermal modalities. Furthermore, a Feature Decoding Module (FDM) that primarily uses three attention mechanisms is meticulously designed, incorporating various attention mechanisms to decode the features. Extensive experiments on the RGBT-CC dataset and the DroneRGBT dataset demonstrate that our proposed MMFFNet has achieved state-of-the-art performance in RGB-T crowd counting. The source code and pretrained models will be released upon acceptance at https://github.com/67zhanan/MMFFNet.
科研通智能强力驱动
Strongly Powered by AbleSci AI