计算机科学
分割
卷积神经网络
人工智能
计算
像素
图像分割
模式识别(心理学)
算法
作者
Wen Luo,Fei Deng,Peifan Jiang,Xiujun Dong,Gulan Zhang
标识
DOI:10.1109/lgrs.2024.3398804
摘要
In recent years, Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) have become the mainstream segmentation methods for high-resolution remote sensing images (HRSIs). CNNs can quickly acquire the correlation between local neighboring pixels through convolutional operations, but it is difficult to establish global contextual relationships, resulting in limited segmentation accuracy. ViTs are able to establish reliable global semantic dependencies through the mechanism of self-attention, but the quadratic computational complexity of self-attention makes the ViTs present high accuracy but low efficiency. Therefore, in this letter, to balance the efficiency and accuracy of HRSIs segmentation, we combine the respective advantages of CNNs and ViTs to propose the FSegNet network. Specifically, we introduce FasterViT and utilize its efficient hierarchical attention to mitigate the surge in self-attention computation due to the high resolution of HRSIs. On this basis, we construct a lightweight decoder based on intensive computation, which achieves fast generation of segmentation results by reshaping and mapping multi-level features. Experiments on the ISPRS Potsdam and Vaihingen datasets show that the proposed FSegNet best balances performance and efficiency. The code is available at https://github.com/Rowan-L/FSegNet.
科研通智能强力驱动
Strongly Powered by AbleSci AI