计算机科学
计算机视觉
目标检测
人工智能
计算机图形学(图像)
对象(语法)
弹丸
遥感
地质学
模式识别(心理学)
化学
有机化学
作者
Tianying Liu,Shuigeng Zhou,Wengen Li,Yichao Zhang,Jihong Guan
标识
DOI:10.1109/tgrs.2025.3550372
摘要
Few-shot object detection (FSOD) has been proposed to solve the problem of insufficient data for training, and it has drawn the attention of the remote sensing community in recent years. A mainstream type of FSOD method is to generate class prototypes based on the limited samples to help the construction of classification decision boundaries. However, these constructed prototypes may be far away from the true class centroids in the few-shot scenario. Recently, the vision-language model (VLM) has shown its powerful ability to align the visual features and text features, which leads to strong zero-shot performance on various downstream computer vision tasks when given only texts. Therefore, in this work, we propose to build class prototypes from text descriptions instead of limited visual instances by leveraging a classical pretrained VLM named CLIP. Concretely, we generate prototypes by feeding the CLIP text encoder with class names and enforcing each positive proposal feature to be close to the corresponding prototype. To accelerate the alignment process, we utilize the CLIP visual encoder as another teacher to achieve visual knowledge distillation. Moreover, we adopt prompt tuning to adapt CLIP to the remote sensing scenario. Extensive experiments on two public FSOD datasets, i.e., DIOR and NWPU VHR-10.v2, demonstrate the effectiveness of our method, which yields competitive results with that of existing approaches.
科研通智能强力驱动
Strongly Powered by AbleSci AI