计算机科学
突出
目标检测
建筑
计算机视觉
变压器
人工智能
模式识别(心理学)
工程类
电气工程
艺术
视觉艺术
电压
作者
Bei Cheng,Zao Liu,Huxiao Tang,Qingwang Wang,Wenhao Chen,Tao Chen,Tao Shen
标识
DOI:10.1109/lgrs.2025.3601083
摘要
The latest remote sensing image saliency detectors primarily rely on RGB information alone. However, spatial and geometric information embedded in depth images is robust to variations in lighting and color. Integrating depth information with RGB images can enhance the spatial structure of objects. In light of this, we innovatively propose a remote sensing image saliency detection model that fuses RGB and depth information, named the multimodal guided transformer architecture (MGTA). Specifically, we first introduce the strong correlated complementary fusion (SCCF) module to explore cross-modal consistency and similarity, maintaining consistency across different modalities while uncovering multidimensional common information. Additionally, the global-local context information interaction (GLCII) module is designed to extract global semantic information and local detail information, effectively utilizing contextual information while reducing the number of parameters. Finally, a cascaded feature-guided decoder (CFGD) is employed to gradually fuse hierarchical decoding features, effectively integrating multi-level data and accurately locating target positions. Extensive experiments demonstrate that our proposed model outperforms 14 state-of-the-art methods. The code and results of our method are available at https://github.com/Zackisliuzao/MGTANet.
科研通智能强力驱动
Strongly Powered by AbleSci AI