计算机科学
情态动词
人工智能
无人机
自然语言处理
图像(数学)
语义学(计算机科学)
图像检索
情报检索
计算机视觉
遗传学
生物
化学
高分子化学
程序设计语言
作者
Jinghao Huang,Yaxiong Chen,Shengwu Xiong,Xiaoqiang Lu
标识
DOI:10.1109/tgrs.2024.3443197
摘要
The cross-modal drone image-text (DIT) retrieval task involves using either text or drone images as queries to retrieve relevant drone images or corresponding text. The primary challenge stems from the diverse and intricate nature of drone images, making effective alignment between image and text challenging. In response, we propose an innovative approach called visual contextual semantic reasoning (VCSR), aimed at precisely aligning information across different modalities. VCSR employs textual cues to guide rich semantic reasoning within the visual context, reducing redundancy in visual information. Furthermore, the method captures drone image information relevant to the text, revealing subtle correspondences between drone image regions and textual content. To enhance visual semantic learning, context region learning (CRL) term and consistency semantic alignment (CSA) terms are introduced for stronger guidance, further intensifying the cross-modal interaction between textual and visual data, resulting in more robust feature representation. Extensive experiments conducted on two self-constructed DIT datasets demonstrate that VCSR outperforms alternative methods in terms of DIT retrieval performance. The codes are accessible at https://github.com/huangjh98/VCSR.
科研通智能强力驱动
Strongly Powered by AbleSci AI