计算机科学
过程(计算)
答疑
背景(考古学)
信息过载
质量(理念)
情报检索
数据科学
生产(经济)
知识管理
噪音(视频)
语义学(计算机科学)
相关性(法律)
万维网
知识图
工业生产
作者
Cong Wang,Shuowen Chai,Tianyi Xu,Muhammad Adil,Tie Qiu
标识
DOI:10.1109/jiot.2026.3652422
摘要
With the increasing adoption of IIoT in industrial production producing massive heterogeneous data, Retrieval-Augmented Generation (RAG) has become a promising approach for industrial knowledge-based Question Answering (QA). However, retrieved top-k documents often contain distracting content that degrades the quality of the generation. Existing research focuses on optimizing retrieval and reranking while overlooking semantic enrichment of useful information and targeted handling of distractions. To address this issue, we propose CP-RAG (Categorize and Process RAG), which categorizes retrieved documents using attention scores and processes them with tailored strategies to mitigate distracting content. Direct-assist documents, which contain concentrated useful information, are enhanced via multi-level semantic optimization to enhance information density. Indirect-assist documents, which carry contextual but distracting elements, are processed through rearrangement with noise mixing to mitigate interference. To maximize LLMs utility, direct-assist documents are placed at both ends of the context window, while indirect-assist documents are positioned centrally. Experimental results show that CP-RAG improves QA accuracy and demonstrates strong practical effectiveness in industrial systems, supporting intelligent decision making over heterogeneous industrial data streams.
科研通智能强力驱动
Strongly Powered by AbleSci AI