计算机科学
人工智能
计算机视觉
图像分割
分割
注释
遥感
模式识别(心理学)
地质学
作者
Yujia Chen,Hao Cui,Guo Zhang,Xue Li,Zhigang Xie,Haifeng Li,Deren Li
标识
DOI:10.1109/tgrs.2024.3523537
摘要
Although significant advances have been made in the semantic segmentation of high-resolution remote sensing (RS) images, obtaining accurate pixelwise annotations remains resource-intensive. We propose SparseFormer, a credible dual-convolutional neural network (CNN) expert-guided Transformer model designed for semantic segmentation using point-level annotations to reduce this annotation burden. SparseFormer comprises three branches, where two CNN branches employ different attention mechanisms to encourage diverse outputs. To enhance the local consistency of pseudolabels, we introduce a pixel-adaptive refinement (PAR) module that dynamically refines CNN output probabilities by incorporating image information during training. A credible assessment is then performed to combine the CNN outputs, producing high-quality pseudolabels that supervise the CNN-Transformer hybrid branch. This hybrid branch integrates global representations with local features, achieving precise segmentation. To further strengthen the CNN branches, we introduce a knowledge distillation strategy that steadily feeds back information from the hybrid branch to CNN branches, mitigating overfitting risks caused by sparse supervision. SparseFormer employs credible assessment to reduce pseudolabel uncertainty, followed by continuous interaction and dynamic information enhancement among the three branches in an end-to-end training process. Extensive experiments on two benchmark datasets demonstrate that SparseFormer significantly outperforms state-of-the-art methods. Our code is available at: https://github.com/Yujia73/SparseFormer.
科研通智能强力驱动
Strongly Powered by AbleSci AI