An adaptive density clustering algorithm for massive data
作者
Keyan Cao,Ibrahim Musa,Jiadi Liu,Yunting Zhang
标识
DOI:10.1109/fskd.2017.8393022
摘要
In this paper, two clustering algorithms are proposed: DBSCAN Entropy-based (DBSCAN) and dynamic clustering algorithm (DBSCAN) to determine the optimal clustering results. The Optimal Number of Clusters ENDBSCAN (OP-ENDBSCAN). ENDBSCAN takes information entropy as the main consideration in clustering, and avoids the traditional DBSCAN algorithm needs to define two parameters of Eps (neighborhood radius) and Minpts (density threshold). At the same time, in order to solve the problem of huge amount of data, a data preprocessing method is proposed. The method divides the data into blocks and divides them into different computer nodes, so as to make full use of the data nodes. Computing power, and improve the efficiency and scalability of the clustering algorithm. OP-ENDBSCAN is an algorithm to determine the optimal number of clustering dynamically and to evaluate the quality of clustering. Based on the analysis of ENDBSCAN, it is found that this algorithm needs to determine the number of clustering by artificially. In order to avoid this problem, OP-ENDBSCAN The effect of anthropogenic parameters on the clustering results was improved and the quality of clustering was improved. Experiments show that both ENDBSCAN and OP-ENDBSCAN can show high efficiency under different data sets and show good clustering results.