计算机科学
可扩展性
钥匙(锁)
答疑
语义数据模型
代表(政治)
智能代理
人工智能
人机交互
智能决策支持系统
期限(时间)
芯(光纤)
组分(热力学)
知识表示与推理
软件工程
多通道交互
语义学(计算机科学)
数据建模
领域(数学)
数据科学
利用
情报检索
作者
Wei Tu,Yuanyuan Li,Man Li,Yuanhao Qiu
摘要
Traditional question-answering systems exhibit significant limitations when processing multimodal and heterogeneous information, restricting their effectiveness in complex scenarios. To address this challenge, this paper proposes an intelligent question-answering system based on a multimodal Retrieval-Augmented Generation (RAG) framework, focusing on key technical challenges such as multi-source heterogeneous data processing and cross-modal semantic alignment. This framework integrates three core technologies: BGE-M3 for text embedding, CLIP for cross-modal representation learning, and BGE-Reranker for precise reordering. Through a multi-stage retrieval mechanism and context-aware generation strategy, it effectively processes multimodal data including text, images, and tables, achieving unified cross-modal semantic modeling and efficient retrieval. At the implementation level, we designed and deployed a highly concurrent, scalable intelligent question-answering platform for university faculty and students, validating its stability and responsiveness at Wuhan University of Science and Technology. Experimental and testing results demonstrate that the system outperforms traditional approaches in retrieval accuracy, response speed, and user interaction satisfaction, fully showcasing the feasibility and advantages of the proposed framework. This research provides a referenceable technical pathway for the engineering implementation of multimodal RAG frameworks in intelligent question-answering systems, laying a foundation for future development of multimodal data-driven intelligent interaction systems.
科研通智能强力驱动
Strongly Powered by AbleSci AI