人工智能
计算机视觉
计算机科学
目标检测
对象(语法)
噪音(视频)
特征(语言学)
模式识别(心理学)
对象类检测
可视化
作者
Jinchuan Wei,Ming Ni,Zhixian Tang,Zhiyong Peng,Shuyao Zhao
摘要
This paper proposes a cue optimization technique for object detection in complex scenarios using a large language model as guidance. By developing a cue optimization algorithm and coupling it closely with the object detection framework, the proposed technique is able to improve object detection in complex scenarios with issues of crowd, occlusions, and multiple objects. The proposed technique takes advantage of semantic understanding properties of the large language model in order to produce initial cues for object detection and uses gradient descent optimization with regularization techniques for iterative optimization of cue parameters. Experimental studies reveal that the optimized technique results in a significant improvement in mean precision, accuracy, recall, F1 measurement, convergence of loss values, error distribution for object detection, and accurate classification of targets. These results validate the usefulness of a novel cue engineering technique for robust object detection with better generalization properties through a novel multimodal visual framework.
科研通智能强力驱动
Strongly Powered by AbleSci AI