选择(遗传算法)
模式识别(心理学)
特征(语言学)
数据挖掘
高维数据聚类
最小冗余特征选择
机器学习
模糊聚类
作者
Salem Alelyani,Jiliang Tang,Huan Liu
出处
期刊:Chapman and Hall/CRC eBooks
[Informa]
日期:2018-09-03
卷期号:: 29-60
被引量:375
标识
DOI:10.1201/9781315373515-2
摘要
Dimensionality reduction techniques can be categorized mainly into feature extraction and feature selection. In the feature extraction approach, features are projected into a new space with lower dimensionality. Feature selection is broadly categorized into four models: filter model, wrapper model, embedded model, and hybrid model. With the existence of a large number of features, learning models tend to overfit and their learning performance degenerates. Feature selection is one of the most used techniques to reduce dimensionality among practitioners. The existence of irrelevant features in the data set may degrade learning quality and consume more memory and computational time that could be saved if these features were removed. However, finding clusters in high-dimensional space is computationally expensive and may degrade the learning performance. Clustering is useful in several machine learning and data mining tasks including image segmentation, information retrieval, pattern recognition, pattern classification, and network analysis.
科研通智能强力驱动
Strongly Powered by AbleSci AI