计算机科学
可执行文件
Python(编程语言)
视觉推理
可视化
人工智能
诱因推理
可扩展性
任务(项目管理)
注释
多通道交互
人机交互
多模态
多模式学习
杠杆(统计)
透明度(行为)
程序设计语言
过程(计算)
自然语言处理
答疑
任务分析
可视对象
脚本语言
自动推理
作者
Zuyi Zhou,Dizhan Xue,Shengsheng Qian
标识
DOI:10.1109/ickg66886.2025.00059
摘要
The Visual Question Answering (VQA) task requires not only accurate answers but also interpretable reasoning processes, particularly in real-world applications where transparency is critical. To reduce annotation and computational costs while maintaining interpretability, the Few-shot Multimodal Explainable VQA (FS-MEVQA) task has been introduced, which aims to generate explanations with limited supervision. In this work, we propose OPeMer (One-shot Prompting and Execution-driven Multimodal Explainable Reasoning), a code-based framework that leverages large language models (LLMs) to generate executable Python programs for multimodal reasoning in a oneshot setting. These programs interact with a lightweight Python API to process visual inputs, capture intermediate reasoning artifacts—such as object crops and spatial relations—and optionally call external tools for open-world visual understanding. The resulting execution traces are serialized and provided to the LLM via a secondary prompt, enabling the generation of coherent multimodal explanations grounded in both visual and textual evidence. Designed without reliance on handcrafted rules or large-scale supervision, OPeMer offers an efficient and extensible approach to explainable multimodal reasoning. Experimental results on the SME dataset demonstrate that OPeMer achieves strong answer accuracy and explanation quality, even when using cost-effective LLMs under limited supervision, suggesting its potential for scalable and interpretable VQA.
科研通智能强力驱动
Strongly Powered by AbleSci AI