人工智能
计算机科学
模式识别(心理学)
分割
图像分割
计算机视觉
特征(语言学)
尺度空间分割
融合
特征提取
传感器融合
点(几何)
图像融合
基于分割的对象分类
作者
Yujia Chen,Hao Cui,Zuohang Wu,Bin Lei,G Zhang,L. Zhang,Chunyang Zhu,Tian Feng,Gui Gao
标识
DOI:10.1109/tgrs.2026.3685134
摘要
Deep learning has made significant progress in the semantic segmentation of remote sensing images. However, its great performance heavily depends on high-quality annotations, which greatly restrict further advancement. Recently, the Segment Anything Model (SAM), enabled by large-scale pre-training and flexible prompt mechanisms, has shown exceptional zero-shot segmentation ability for any object. Hence, we devise RSPoint-SAM, a SAM-based point-supervised semantic segmentation model that can achieve excellent results using point annotations. RSPoint-SAM has two main branches: convolutional neural network (CNN) and SAM-driven fusion. The CNN branch uses EfficientNet as its feature extraction backbone and employs a cross-level feature fusion strategy to combine multi-level features, effectively capturing local contextual dependencies under sparse supervision. Next, high-confidence pseudo-labels are produced by combining the first two outputs of the CNN branch with a credible assessment mechanism. The SAM-driven fusion branch then integrates EfficientViT and EfficientNet to utilize the ability of both global context modeling of the vision transformer and the structured boundary perception of SAM, thereby capturing object-level semantics under limited pseudo-label supervision. Moreover, we develop a knowledge dynamic interaction mechanism to feed global semantic and boundary information from the SAM-driven fusion branch back into the CNN branch, reducing overfitting and further reinforcing pseudo-label quality. Experiments on two benchmark datasets show that RSPoint-SAM outperforms state-of-the-art weakly supervised methods. The source code will be available at: https://github.com/Yujia73/RSPoint-SAM.
科研通智能强力驱动
Strongly Powered by AbleSci AI