计算机科学
弹丸
人工智能
一次性
人机交互
计算机视觉
自然语言处理
工程类
机械工程
有机化学
化学
作者
Jishnu Jaykumar P,Kamalesh Palanisamy,Yu-Wei Chao,Xinya Du,Xiang Yu
标识
DOI:10.1109/iros58592.2024.10801660
摘要
We propose a novel framework for few-shot learning by leveraging large-scale vision-language models such as CLIP [1]. Motivated by unimodal prototypical networks for few-shot learning, we introduce Proto-CLIP which utilizes image prototypes and text prototypes for few-shot learning. Specifically, Proto-CLIP adapts the image and text encoder embeddings from CLIP in a joint fashion using few-shot examples. The embeddings from the two encoders are used to compute the respective prototypes of image classes for classification. During adaptation, we propose aligning the image and text prototypes of the corresponding classes. Such alignment is beneficial for few-shot classification due to the reinforced contributions from both types of prototypes. Proto-CLIP has both training-free and fine-tuned variants. We demonstrate the effectiveness of our method by conducting experiments on benchmark datasets for few-shot learning, as well as in the real world for robot perception1.
科研通智能强力驱动
Strongly Powered by AbleSci AI