计算机科学
图形
图形数据库
加速
数据管理
数据检索
数据挖掘
计算机数据存储
方案(数学)
解码方法
关系数据库
理论计算机科学
并行计算
算法
操作系统
数学
数学分析
作者
Xue Li,Weibin Zeng,Zhibin Wang,Diwen Zhu,Jingbo Xu,Wenyuan Yu,Jingren Zhou
标识
DOI:10.48550/arxiv.2312.09577
摘要
Data lakes, increasingly adopted for their ability to store and analyze diverse types of data, commonly use columnar storage formats like Parquet and ORC for handling relational tables. However, these traditional setups fall short when it comes to efficiently managing graph data, particularly those conforming to the Labeled Property Graph (LPG) model. To address this gap, this paper introduces GraphAr, a specialized storage scheme designed to enhance existing data lakes for efficient graph data management. Leveraging the strengths of Parquet, GraphAr captures LPG semantics precisely and facilitates graph-specific operations such as neighbor retrieval and label filtering. Through innovative data organization, encoding, and decoding techniques, GraphAr dramatically improves performance. Our evaluations reveal that GraphAr outperforms conventional Parquet and Acero-based methods, achieving an average speedup of 4452x for neighbor retrieval, 14.8x for label filtering, and 29.5x for end-to-end workloads. These findings highlight GraphAr's potential to extend the utility of data lakes by enabling efficient graph data management.
科研通智能强力驱动
Strongly Powered by AbleSci AI