A RAG data pipeline transforming heterogeneous data into AI-ready format for autonomous building performance discovery

管道(软件) 计算机科学 自动化 原始数据 数据挖掘 过程(计算) 数据库 数据质量 语义异质性 情报检索 数据建模 基线(sea) 语义数据模型 数据提取 体积热力学 数据处理 信息抽取 数据管理 数据集成 数据检索 数据迁移 质量(理念) 非结构化数据 数据结构 数据映射 语义映射 数据整理 数据存取
作者
Han Li,Ana Comesana,Christopher Weyandt,Tianzhen Hong
出处
期刊:Advances in applied energy [Elsevier BV]
卷期号:21: 100261-100261 被引量:5
标识
DOI:10.1016/j.adapen.2025.100261
摘要

• Introduces a novel RAG data pipeline that integrates domain-adapted extraction to convert building documents into accurate, provenance-grounded answers • Converts unstructured multimodal inputs (images, text, tables) into AI-ready, provenance-linked chunks while preserving spatial and hierarchical relationships • Delivers approximately 30% improvement in coverage and structural preservation with domain-adapted RAG processing • Demonstrates potential integration with MCP tools to automatically generate semantic data models, reducing effort from days to minutes. • Substantially reduces data-preparation overhead, enabling accessible building information for operations, auditing, and analytics. . Despite a growing volume of data being collected from buildings, a small portion of the data has been analyzed to provide actionable insights, mainly due to the labor intensive and error prone process of integrating and understanding the heterogeneous data with varying levels of quality and resolutions. This study introduces a novel domain-adapted Retrieval Augmented Generation (RAG) based data pipeline designed for efficient information retrieval and management from heterogeneous building data sources. The proposed data pipeline leverages Large Language Models (LLMs) and domain-specific processing to transform unstructured data (including images, tables, and plain text) into semantically rich, AI-ready representations. Evaluation results show that the domain-adapted processing techniques achieved approximately 30% improvement in both coverage and structural preservation for documents like images and tables compared to generic baseline methods employing generic prompts and minimal preprocessing. This automation capability transforms raw building artifacts into AI-ready searchable knowledge bases. A critical application demonstrated is the automation of semantic data model creation, reducing the manual effort from potentially days to minutes. This domain-adapted RAG pipeline significantly addresses persistent data challenges in the building sector, increasing the accessibility and utility of building information for diverse stakeholders and applications.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
Flz应助猪八戒采纳,获得10
1秒前
不想动欸完成签到,获得积分10
2秒前
ya0701发布了新的文献求助10
3秒前
3秒前
涂涂发布了新的文献求助30
4秒前
5秒前
5秒前
haonanchen发布了新的文献求助10
6秒前
7秒前
CuZn完成签到,获得积分10
9秒前
王晨昕发布了新的文献求助10
10秒前
10秒前
田様应助seven采纳,获得10
10秒前
YRY完成签到,获得积分10
11秒前
Luuu完成签到 ,获得积分10
11秒前
ll完成签到 ,获得积分10
13秒前
14秒前
14秒前
在水一方应助筚路蓝缕采纳,获得10
15秒前
科研通AI6.4应助dffh采纳,获得10
17秒前
机灵的团完成签到,获得积分20
17秒前
正直的手链完成签到,获得积分10
17秒前
xiaohuiben发布了新的文献求助10
17秒前
王加通完成签到,获得积分10
17秒前
上官若男应助拾陆采纳,获得10
17秒前
19秒前
19秒前
20秒前
20秒前
猩猩完成签到,获得积分10
21秒前
21秒前
Liutingli应助内向的含之采纳,获得10
21秒前
丘比特应助幸福遥采纳,获得10
21秒前
负数发布了新的文献求助10
21秒前
酷波er应助梓楠采纳,获得10
22秒前
满意宛筠发布了新的文献求助10
22秒前
涂涂完成签到,获得积分20
22秒前
打打应助斑马123采纳,获得10
24秒前
一只东北鸟完成签到 ,获得积分10
25秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
HYDROLYSE ACIDE DE QUELQUES DIOXASPIROCYCLANES 1314
Essentials of Carbohydrate Chemistry and Biochemistry, 4th Edition 800
Navigating Normative Orders. Interdisciplinary Perspectives 800
1 Peter and Christ's Descent to the Dead in Its Early Christian Reception 700
Organizational Behavior 510
Management and the Arts 510
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7747950
求助须知:如何正确求助?哪些是违规求助? 9296180
关于积分的说明 20233931
捐赠科研通 7329325
什么是DOI,文献DOI怎么找? 3308744
关于科研通互助平台的介绍 2460530
邀请新用户注册赠送积分活动 2320713