计算机科学
人工智能
编码器
分割
卷积神经网络
变压器
特征提取
模式识别(心理学)
图像分割
特征学习
计算机视觉
物理
量子力学
电压
操作系统
作者
Tao Xiao,Yikun Liu,Yuwen Huang,Mingsong Li,Gongping Yang
标识
DOI:10.1109/tgrs.2023.3256064
摘要
Semantic segmentation is an extremely challenging task in high-resolution remote sensing (HRRS) images as objects have complex spatial layouts and enormous variations in appearance. Convolutional neural networks (CNNs) have excellent ability to extract local features and have been widely applied as the feature extractor for various vision tasks. However, due to the inherent inductive bias of convolution operation, CNNs inevitably have limitations in modeling long-range dependencies. Transformer can capture global representations well, but unfortunately ignores the details of local features and has high computational and spatial complexity in processing high-resolution feature maps. In this paper, we propose a novel hybrid architecture for HRRS image segmentation, termed EMRT, to exploit the advantages of convolution operations and Transformer to enhance multi-scale representation learning. We incorporate the deformable self-attention mechanism in the Transformer to automatically adjust the receptive field, and design an encoder-decoder architecture accordingly to achieve efficient context modeling. Specifically, the CNN is constructed to extract feature representations. In the encoder, local features and global representations at different resolutions are extracted by the CNN and Transformer, respectively, and fused in an interactive manner. Moreover, a separate spatial branch is designed to extract multi-scale contextual information as queries, and global dependencies between features at different scales are efficiently established by the decoder. Extensive experiments on three public remote sensing datasets demonstrate the superiority of EMRT and indicate that the overall performance of our method outperforms state-of-the-art methods. Code is available at https://github.com/peach-xiao/EMRT.
科研通智能强力驱动
Strongly Powered by AbleSci AI