场景图
计算机科学
人工智能
图形
动作识别
对象(语法)
答疑
计算机视觉
任务(项目管理)
自然语言处理
可视化
任务分析
视觉语言
知识图
图论
自然语言
语言模型
作者
Weixin Chen,Yongyong Chen,Shichao Kan
标识
DOI:10.1109/icme59968.2025.11210017
摘要
Scene graph generation (SGG) is pivotal for acquiring valuable knowledge in visual scene understanding, making it crucial for tasks such as visual question answering and visual reasoning. In recent times, multimodal large language models (MLLMs) have demonstrated remarkable proficiency in object grounding and recognition. Nevertheless, constructing the scene graph directly poses a challenging task for MLLMs due to the intricate nature of predicting relationships between objects and the action states of objects. Simultaneously, reliance solely on the large language model (LLM) for achieving region understanding in complex scenes makes MLLMs susceptible to hallucinations, leading to potential errors in determining the coordinates of objects. To tackle these challenges, we introduce a large vision-language model (LVLM) within the Shikra framework for scene graph generation. Our approach involves the creation of an instruction-following SGG dataset for the fine-tuning of the LVLM. After SGG, we recognize that the scene graph generated by the LVLM can assist the LLM in answering visual questions. Thus, we evaluate LVLM-based question-and-answering models by leveraging the scene graph as a rationale, introducing a concept termed Scene Graph Chain of Thought (SGCoT). The proposed SGG method is rigorously evaluated through both quantitative and qualitative experiments on closed-set and open-set SGG tasks, affirming its effectiveness. Moreover, using the scene graph generated by fine-tuning LVLM as a rationale in the chain of thoughts results in a competitive performance on several vision-language compositional benchmarks.
科研通智能强力驱动
Strongly Powered by AbleSci AI