遥感
计算机科学
背景(考古学)
分割
计算机视觉
人工智能
图像分割
最小边界框
特征(语言学)
GSM演进的增强数据速率
编码器
遥感应用
航空影像
空间语境意识
目标检测
匹配(统计)
任务(项目管理)
水准点(测量)
适应(眼睛)
跳跃式监视
边缘检测
特征提取
忠诚
详细程度
灵活性(工程)
尺度空间分割
作者
Musarat Hussain,Ji Peng Huang,Xiankui Liu,Yulin Duan,Hongyan Wu
标识
DOI:10.1109/jstars.2025.3649081
摘要
Accurate building segmentation in remote sensing images is crucial for applications like disaster assessment, 3D urban modeling, and monitoring urban transformations. However, this task presents significant challenges due to the vast geographical coverage, dense building clusters, and the complexity of building contours, roof geometries, and surrounding environments. While the Segment Anything Model (SAM) offers a promising solution for extracting building masks in remote sensing images, its reliance on interactive input cues, difficulty in capturing fine edge details, and inability to integrate global semantic context with local fine-grained visual features often result in poor boundary detection and fragmented masks, limiting its effectiveness in fully automated, end-to-end building segmentation. To address these limitations, we propose YOLOSAM, a YOLO-guided adaptation of SAM designed for precise and automated building segmentation. Our framework introduces three lightweight yet effective innovations: (i) an Automatic Prompt Generator, based on YOLOv8, that automatically produces bounding box prompts to eliminate manual input; (ii) a High-Quality Token (HQ-Token) that improves edge fidelity and mask coherence by refining SAM's decoder representations; and (iii) a Global-Local Feature Fusion module, which enhances segmentation quality by fusing semantic context from deeper layers with fine edge details from earlier stages of SAM's frozen architecture. Importantly, our method preserves SAM's pre-trained generalization ability by freezing the original encoder and decoder while training only the lightweight modules. Experimental results demonstrate a significant improvement in segmentation accuracy, with mIoU increasing to 76.7% on the WHU building segmentation dataset, 69.1% on the Vaihingen building dataset, and 73.2% on the Inria Aerial Image Labeling dataset, compared to SAM's “segment everything” mode. Our model also significantly outperforms both classical deep learning baselines and other SAM-based frameworks.
科研通智能强力驱动
Strongly Powered by AbleSci AI