计算机科学
图像检索
集合(抽象数据类型)
人工智能
图像(数学)
简单(哲学)
情报检索
计算机视觉
模式识别(心理学)
数据集
可视化
图像处理
机器学习
数据挖掘
基础(证据)
搜索引擎索引
图像自动标注
光学(聚焦)
视觉文字
作者
Guanqi Zhan,Yuanpei Liu,Kai Han,Weidi Xie,Andrew Zisserman
标识
DOI:10.1109/cbmi66578.2025.11339290
摘要
The objective in this paper is to improve the per-formance of text-to-image retrieval. To this end, we introduce a new framework that can boost the performance of large-scale pre-trained vision-language models, so that they can be used for text-to-image re-ranking. The approach, Enhanced Language-Image Pre-training (ELIP), uses the text query, via a simple MLP mapping network, to predict a set of visual prompts to condition the ViT image encoding. ELIP can easily be applied to the commonly used CLIP, SigLIP and BLIP-2 networks. On the evaluation side, we set up two new out-of-distribution (OOD) benchmarks, Occluded COCO and ImageNet-R, to assess the zero-shot generalisation of the models to different domains. The results demonstrate that ELIP significantly boosts CLIP/SigLIP/SigLIP-2 text-to-image retrieval performance and outperforms BLIP-2 on several benchmarks, as well as providing an easy means to adapt to OOD datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI