高光谱成像
激光雷达
计算机科学
遥感
土地覆盖
深度学习
测距
土地利用
人工智能
变压器
机器学习
地理
电信
物理
工程类
土木工程
量子力学
电压
作者
Swalpa Kumar Roy,Atri Sukul,Ali Jamali,Juan M. Haut,Pedram Ghamisi
标识
DOI:10.1109/tgrs.2024.3374324
摘要
The successes of attention-driven deep models like the Vision Transformer (ViT) have sparked interest in cross-domain exploration. However, current transformer-based techniques in remote sensing primarily focus on single-modal data, limiting their potential to exploit the growing array of multimodal Earth observation data fully. Enhancing these models for multimodal integration is crucial for comprehensive remote sensing applications. To achieve this, we extend the traditional self-attention mechanism by introducing Cross Hyperspectral and LiDAR (Cross-HL) attention. We present a novel multimodal deep learning framework that effectively fuses remote sensing (RS) data, intending to improve land use and land cover (LULC) recognition. To enhance the accurate exchange of information across different modalities, we fuse their patch projections using the Cross-HL self-attention module. In this process, LiDAR patch tokens serve as queries ( Q ), while keys ( K ) and values ( V ) are derived from HS patch tokens. To demonstrate the superiority of Cross-HL in the proposed multimodal deep learning framework, we conducted extensive experiments on three multimodal RS benchmark datasets: Houston, Trento, and MUUFL. These datasets contain hyperspectral and light detection and ranging (LiDAR) data. The source code for Cross-HL will be made available publicly at https://github.com/AtriSukul1508/Cross-HL.
科研通智能强力驱动
Strongly Powered by AbleSci AI