情态动词
计算机科学
光学(聚焦)
上下文图像分类
特征(语言学)
人工智能
匹配(统计)
语义学(计算机科学)
特征向量
特征提取
图像(数学)
模式识别(心理学)
计算机视觉
数学
语言学
哲学
物理
光学
化学
高分子化学
统计
程序设计语言
作者
Chun Liu,Suqiang Ma,Zheng Li,Wei Yang,Zhigang Han
标识
DOI:10.1109/lgrs.2024.3368344
摘要
The task of zero-shot classification of image scenes is to recognize the image scenes that are not seen in the training stage. To address the zero-shot image scene classification problem, the cross-modal feature alignment methods have been proposed in recent years. These methods mainly focus on matching the visual features of each image scene with their corresponding semantic descriptors in the latent space. Less attention has been paid to the contrastive relationships between different image scenes and different semantic descriptors. In this work, we propose a multi-level feature alignment method by mining the contrastive relations between cross-modal features for zero-shot classification of remote sensing image scenes. While promoting the single-instance level positive alignment between each image scene with their corresponding semantic descriptors, the proposed method learns to keep the visual and semantic features of different classes in the latent space apart from each other. Extensive experiments have shown that the proposed method has better performance for zero-shot remote sensing image scene classification. All the code and data are available at github https://github.com/masuqiang/MCFA-Pytorch.
科研通智能强力驱动
Strongly Powered by AbleSci AI