答疑
计算机科学
视觉推理
人工智能
嵌入
变压器
自然语言处理
情报检索
机器学习
物理
量子力学
电压
作者
Swati Srivastava,Himanshu Sharma
标识
DOI:10.1117/1.jei.33.6.063052
摘要
The proposed work for chart question answering (CQA) addresses crucial aspects, particularly those associated with questions requiring complex reasoning and visual references to charts and those related to optical character recognition (OCR) noise. Existing methods focused mainly on image features, neglecting the unique characteristics of the visual representation, which consequently not able to answer complex reasoning questions. Introducing structural embedding, question embedding, and a transformer-based relational understanding of the visualizations enable the proposed model to answer complex reasoning questions. The incorporation of the reading component and visual component using BERT encoding leverages both textual and visual information, which overcomes the OCR noise-related issue. The comprehensive ablation study systematically evaluates the different components of our model. Moreover, experiments are conducted on five datasets: DVQA, FigureQA, PlotQA, LEAFQA++, and ChartQA, which makes the assessment robust and generalized. The overall accuracy of our model is increased by 1.15% compared with STL-CQA on the test familiar of DVQA, 0.55% compared with PReFIL on FigureQA, 8.03% compared with Plot-QA, 2.28% compared with STL-CQA test-familiar split of LEAFQA++, and 0.40% compared with VisionTaPas on the test-set of ChartQA, which shows the effectiveness and superiority of the proposed model over existing methods on these datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI