分布式文件系统
计算机科学
复制(统计)
文件系统
分布式数据存储
重复数据消除
操作系统
容错
数据库
集合(抽象数据类型)
分布式数据库
分布式计算
程序设计语言
数学
统计
作者
D. Veeraiah,J. Nageswara Rao
标识
DOI:10.1109/icict48043.2020.9112567
摘要
HDFS [Hadoop Distributed File System] a part of Apache Hadoop to store large data set consistently. HDFS is used for process Massive-Scale Data in parallel and it ensures accessibility of facts by replicating data to different nodes. Still, the repetition policy of HDFS doesn't think about the name of knowledge. The recognition of the files tends to alter over time. Hence, maintaining a fixed replication issue can affect the storage efficiency of HDFS. An Efficient Data Duplication System Based on HDFS, is proposed which consider the reputations of the records set aside in HDFS before replication. The proposed technique successfully reduces storage consumption by up to 45% without moving the accessibility and fault recognition in HDFS. □
科研通智能强力驱动
Strongly Powered by AbleSci AI