分割
表达式(计算机科学)
理解力
计算机科学
人工智能
遥感
计算机视觉
自然语言处理
地理
程序设计语言
作者
Xiaoqiang Lu,Long Sun,Lingling Li,Licheng Jiao,Yuting Yang,Zhongjian Huang,Jinming Chai,Xu Liu,Fang Liu,Wenping Ma,Shuyuan Yang
标识
DOI:10.1109/mgrs.2025.3574685
摘要
Understanding and interpreting a specific object from large-scale remote sensing (RS) scenes provide basic support for various practical applications. To achieve it, visual grounding (VG) and referring image segmentation (RIS) are two main techniques that aim to localize and segment the referred object given a free-form linguistic expression. Currently, most works of VG and RIS focus on natural images, with only a few generalizing to RS images and further developing remote sensing visual grounding (RSVG) and referring remote sensing image segmentation (RRSIS). However, RSVG and RRSIS are designed to solve separate tasks, ignoring the benefits of jointly learning localization and segmentation. In this work, we introduce the task of referring remote sensing expression comprehension and segmentation (RRSECS) to explore the potential of multi-task learning in vision-language understanding. Specifically, we construct the first benchmark for this task, namely RefDIOR, which contains image-expression-box-mask quadruplets for training and evaluating different models, enabling us to advance the research of RRSECS. Then, we benchmark extensive methods across VG, RSVG, RIS, and RRSIS on RefDIOR, and give insightful analyses of their performances and limitations. Finally, we propose a novel cross-task collaborative Transformer (CCFormer) to accomplish language-guided end-to-end localization and segmentation, serving as a strong baseline for RRSECS. CCFormer consists of the multi-scale cross-modal fusion module to obtain fine-grained aligned vision-language features, the language-aware gated decoupling module to assign discriminative multi-modal features for each task, and the cross-task collaborative loss to refine multiple outputs in a co-rectify manner. Experimental results on RefDIOR demonstrate the effectiveness and generalization of CCFormer in addressing the challenges of RSVG and RRSIS. The code and dataset will be publicly accessible at https://github.com/IPIU-XDU/RSFM.
科研通智能强力驱动
Strongly Powered by AbleSci AI