遥感
计算机科学
分割
基础(证据)
图像分割
人工智能
遥感应用
计算机视觉
光学成像
图像(数学)
语义学(计算机科学)
数据建模
作者
Bowei Ye,Tao Shao,Fei Su,Zhicheng Zhao
标识
DOI:10.1109/tgrs.2026.3677380
摘要
Deep learning algorithms have driven substantial progress in remote sensing semantic segmentation. However, conventional approaches typically rely on predefined semantic categories, necessitating costly data annotation and model retraining when new classes are introduced. While large-scale vision-language models, such as CLIP, enable segmentation of arbitrary class with natural language guidance, their limited localization capability poses challenges for dense prediction tasks. This study investigates the potential of CLIP for semantic segmentation of optical remote sensing images and proposes LD-Seg, a novel training-free framework that optimizes localization ability and visual-textual alignment through dual representation refinement. Through systematic analysis, we identify anomalous “singularity” tokens in the CLIP visual encoder that aggregate global contextual information while disproportionately attracting attention from local patch tokens, thereby degrading spatial discriminability. To address this, we introduce a singularity feature repair (SFR) strategy that mitigates feature distortion by recalibrating these tokens. Furthermore, we develop a hierarchical semantic expansion (HSE) method to generate precise hierarchical text descriptions, enhancing cross-modal alignment. The SFR and HSE strategies complement existing methods, providing further improvements. Extensive experiments demonstrate that LDSeg achieves state-of-the-art performance, delivering mIoU improvements of 0.95% to 7.84% over existing methods. Code is available at https://github.com/yeelinsen/LDSeg.
科研通智能强力驱动
Strongly Powered by AbleSci AI