计算机科学
编码器
语义学(计算机科学)
人工智能
分割
计算机视觉
代表(政治)
图像分割
对偶(语法数字)
一致性(知识库)
图像(数学)
编码(内存)
图像融合
模式识别(心理学)
编码(集合论)
自然语言
空间分析
语义鸿沟
遥感
中间语言
融合
可视化
传感器融合
目标检测
隐藏字幕
图像分辨率
作者
Jingwen Zhang,Lingling Li,Licheng Jiao,Xu Liu,Fang Liu,Wenping Ma,Shuyuan Yang
标识
DOI:10.1109/tgrs.2026.3651598
摘要
Referring Remote Sensing Image Segmentation (RRSIS) aims to segment target objects in remote sensing images using natural language descriptions. Existing methods employing a single backbone and sequential fusion struggle to capture fine-grained semantics in remote sensing data due to their complex and multi-scale nature. Moreover, vision models pretrained on natural images often fail to generalize to remote sensing images due to domain-specific spatial and semantic discrepancies. To address these issues, we propose an innovative Multiscale Vision-Text Collaborative Dual-Encoder network, named MCD-Net. We first introduce a frozen SAM encoder as a structure-aware auxiliary branch to inject general-purpose spatial priors, enhancing the Swin Transformer’s ability to model fine-grained geometry and object-level information. To improve semantic consistency across scales and better align with referring expressions, we propose a Multi-scale Frequency-aware Alignment module that decomposes visual features into frequency components and modulates them via cross-scale textual attention. An Adaptive Deep Fusion module bridges the representation gap between dual visual branches while preserving spatial-semantic coherence. Extensive experiments on public RRSIS benchmarks demonstrate that our method outperforms state-of-the-art approaches, particularly in scenes with dense objects and ambiguous expressions. Our code will be released upon publication.
科研通智能强力驱动
Strongly Powered by AbleSci AI