变压器
计算机科学
公制(单位)
人工智能
任务(项目管理)
相似性(几何)
机器学习
领域(数学)
钥匙(锁)
数据挖掘
数学
图像(数学)
工程类
电气工程
纯数学
系统工程
电压
计算机安全
运营管理
标识
DOI:10.1109/icicml57342.2022.10009711
摘要
With the advent of the information age, VQA has become a key research direction cross the field of CV and NLP. Based on a large number of real scene images on the KAGGLE platform, this paper applies the Transformers model to the CV field, and combines the Transformers based NLP algorithm to establish a VQA system. The experimental results show that our model can give accurate answers in a simple and orderly scenario, and there is a certain deviation between the generated results and the real answers in a chaotic scenario, which verifies the effectiveness of our model. Wup similarity is selected as a metric in this paper. The results show that the RoBERTa-ViT model has the most outstanding performance in VQA task, with the metric value of 0.351. Surprisingly, the BERT-SwinT model performs relatively poorly. which may be because the SwinT model is more complex than the VIT model and does not show its own advantages in small data set task.
科研通智能强力驱动
Strongly Powered by AbleSci AI