编码器
对偶(语法数字)
计算机科学
融合
人工智能
计算机视觉
图像(数学)
图像融合
算法
模式识别(心理学)
语言学
操作系统
文学类
哲学
艺术
摘要
Many sectors are challenged by how to effectively represent knowledge in files that contain multiple images closely related to text, and how to make models understand the relationship between images and text. Contrastive Language-Image Pre-training (CLIP) and Bootstrapping Language-Image Pre-training (BLIP) acquire the capability of understanding the image-text relationship through large-scale model pre-training. CLIP not only considers images and their related text but also contrasts images with massive irrelevant text, to improve its capability of generalizing the relationship between images and related text. BLIP enhances its understanding of complex image-text relationships by pre-training and fine-tuning matched image-text pairs. This paper presents an image-text fusion algorithm based on CLIP and BLIP, which gives an accurate and consistent picture of the image-text relevance by fully using CLIP's image-text generalization capacity and BLIP's capacity of understanding the complex image-text relationship.
科研通智能强力驱动
Strongly Powered by AbleSci AI