计算机科学
分割
编码器
人工智能
计算机视觉
图像分割
遥感
变压器
尺度空间分割
图像分辨率
基于分割的对象分类
遥感应用
高分辨率
图像(数学)
目标检测
解码方法
路径(计算)
对象(语法)
模式识别(心理学)
作者
Siting Xiong,Linfeng Wu,Bochen Zhang,Dejin Zhang,Yu Tao,Yuzhi Tang
标识
DOI:10.1109/tgrs.2026.3655448
摘要
The groundbreaking segment anything model (SAM), built on a vision transformer (ViT) design with millions of parameters and trained on the large SA-1B dataset, acts as a vision foundation model that can be used for various segmentation tasks. However, this model cannot be directly applied to the semantic segmentation of remote sensing images, as it tends to over-segment objects rather than preserving their semantics. Furthermore, it does not align well with object boundaries, which are usually required for the high-accuracy segmentation of high-resolution remote sensing images. To address this issue, we adjust the image encoder and mask decoder of SAM and propose an HL-SAM-Seg network. The image encoder is extensively adjusted to comprise an adapter, a high-resolution (high-res) path, and a low-resolution (low-res) path. The latter two paths extract and update the high-res and low-res features, which are merged in the mask decoder to perform multi-class segmentation. Specifically, we design a highlow- resolution cross-attention (HL) module inserted into the transformer blocks of the low-res path to align and update the low-res and high-res features. Experimental results on the ISPRS Vaihingen, ISPRS Potsdam, and FloodNet datasets show that the proposed HL-SAM-Seg outperformed conventional state-of-theart semantic segmentation algorithms overall, with underperformance in some small sample categories. Moreover, it has fewer trainable parameters, suggesting the potential for leveraging SAM for remote sensing image segmentation.
科研通智能强力驱动
Strongly Powered by AbleSci AI