计算机科学
目标检测
模拟退火
人工智能
计算机视觉
模式识别(心理学)
机器学习
作者
Qihao Chen,Jiaxuan Zheng
标识
DOI:10.1109/iceace63551.2024.10898935
摘要
With the development of multimodal learning technology, the demand for combining text information to achieve more accurate object detection tasks is growing. Traditional object detection models usually rely on specific visual features. However, with the increase in task complexity, pure visual information is often not enough to meet actual needs. Therefore, it has become a trend to introduce semantic information of natural language. This paper proposes a multimodal object detection model that combines the CLIP model with the YOLO model and optimizes it using the simulated annealing algorithm. First, YOLO is used for preliminary object detection to obtain bounding boxes and category predictions. Then, CLIP, which has been fine-tuned using the Chinese dataset, is used to enhance the features of each detected object region, calculate the similarity between the image and the text description, and set and optimize the image-text matching threshold to 86% using the simulated annealing algorithm to achieve more accurate reclassification. We introduce the contrast loss of CLIP into the YOLO loss function to form a joint optimization framework, which enables the model to improve the semantic understanding ability while retaining the efficiency of YOLO. Experimental results show that the combination of CLIP and YOLO significantly improves the performance of object detection, especially in complex scenes and multimodal information understanding. This research provides new ideas and methods for the future integration of computer vision and natural language processing.
科研通智能强力驱动
Strongly Powered by AbleSci AI