聚类分析
计算机科学
数据挖掘
人工智能
离群值
启发式
高维数据聚类
线性判别分析
贝叶斯概率
多元统计
机器学习
共识聚类
模式识别(心理学)
相关聚类
CURE数据聚类算法
作者
Chris Fraley,Adrian E. Raftery
出处
期刊:
日期:2002-06-01
卷期号:97 (458): 611-631
被引量:4164
标识
DOI:10.1198/016214502760047131
摘要
Cluster analysis is the automated search for groups of related observations in a dataset. Most clustering done in practice is based largely on heuristic but intuitively reasonable procedures, and most clustering methods available in commercial software are also of this type. However, there is little systematic guidance associated with these methods for solving important practical questions that arise in cluster analysis, such as how many clusters are there, which clustering method should be used, and how should outliers be handled. We review a general methodology for model-based clustering that provides a principled statistical approach to these issues. We also show that this can be useful for other problems in multivariate analysis, such as discriminant analysis and multivariate density estimation. We give examples from medical diagnosis, minefield detection, cluster recovery from noisy data, and spatial density estimation. Finally, we mention limitations of the methodology and discuss recent developments in model-based clustering for non-Gaussian data, high-dimensional datasets, large datasets, and Bayesian estimation.
科研通智能强力驱动
Strongly Powered by AbleSci AI