隐藏字幕
计算机科学
特征提取
遥感
特征(语言学)
人工智能
信息抽取
变更检测
图像(数学)
情报检索
语义特征
计算机视觉
地质学
语言学
哲学
作者
Renlong Hang,Jinyu Luo,Hui Lin,Qingshan Liu
标识
DOI:10.1109/tgrs.2025.3595422
摘要
Remote sensing image change captioning (RSICC) aims to generate sentence descriptions about land cover changes in bitemporal images. The effective acquisition of semantic-level change information is critical for this task. However, due to the effects of illumination interference, appearance similarities and scale differences between different objects, it is difficult to accurately extract change information from bitemporal images. In this article, we attempt to take advantage of the high-level semantic information inherent in text and propose a text-augmented semantic feature extraction and difference information learning model for RSICC. Specifically, we first pre-define some text prompts for each remote sensing image and use the contrastive language-image pretraining (CLIP) model to select the most suitable text descriptions for them. Then, we adopt a refined segment anything model (SAM) to learn fine-grained visual features from each image, which is further enhanced via a designed selective text-image fusion (STIF) module. After that, to extract the semantic differences between bitemporal images, we propose a text-guided difference capture (TGDC) module capable of extracting multiscale difference information under the guidance of text differences between different-time images. Finally, a transformer-based caption generator is applied to generate sentence descriptions from the extracted difference information. In order to test the performance of our proposed model, we conduct comprehensive experiments on two widely used RSICC datasets, including LEVIR-CC and Dubai-CC. The experimental results show that our proposed model is able to outperform several state-of-the-art models, which validates the effectiveness of it. The codes of our proposed model will be released at https://github.com/Richardkimyo/TACC.
科研通智能强力驱动
Strongly Powered by AbleSci AI