管道(软件)
计算机科学
自动化
原始数据
数据挖掘
过程(计算)
数据库
数据质量
语义异质性
情报检索
数据建模
基线(sea)
语义数据模型
数据提取
体积热力学
数据处理
信息抽取
数据管理
数据集成
数据检索
数据迁移
质量(理念)
非结构化数据
数据结构
数据映射
语义映射
数据整理
数据存取
作者
Han Li,Ana Comesana,Christopher Weyandt,Tianzhen Hong
标识
DOI:10.1016/j.adapen.2025.100261
摘要
• Introduces a novel RAG data pipeline that integrates domain-adapted extraction to convert building documents into accurate, provenance-grounded answers • Converts unstructured multimodal inputs (images, text, tables) into AI-ready, provenance-linked chunks while preserving spatial and hierarchical relationships • Delivers approximately 30% improvement in coverage and structural preservation with domain-adapted RAG processing • Demonstrates potential integration with MCP tools to automatically generate semantic data models, reducing effort from days to minutes. • Substantially reduces data-preparation overhead, enabling accessible building information for operations, auditing, and analytics. . Despite a growing volume of data being collected from buildings, a small portion of the data has been analyzed to provide actionable insights, mainly due to the labor intensive and error prone process of integrating and understanding the heterogeneous data with varying levels of quality and resolutions. This study introduces a novel domain-adapted Retrieval Augmented Generation (RAG) based data pipeline designed for efficient information retrieval and management from heterogeneous building data sources. The proposed data pipeline leverages Large Language Models (LLMs) and domain-specific processing to transform unstructured data (including images, tables, and plain text) into semantically rich, AI-ready representations. Evaluation results show that the domain-adapted processing techniques achieved approximately 30% improvement in both coverage and structural preservation for documents like images and tables compared to generic baseline methods employing generic prompts and minimal preprocessing. This automation capability transforms raw building artifacts into AI-ready searchable knowledge bases. A critical application demonstrated is the automation of semantic data model creation, reducing the manual effort from potentially days to minutes. This domain-adapted RAG pipeline significantly addresses persistent data challenges in the building sector, increasing the accessibility and utility of building information for diverse stakeholders and applications.
科研通智能强力驱动
Strongly Powered by AbleSci AI