自编码
聚类分析
计算机科学
嵌入
人工智能
地点
代表(政治)
特征学习
深度学习
理论计算机科学
机器学习
模式识别(心理学)
数据挖掘
政治学
哲学
政治
法学
语言学
作者
Bassoma Diallo,Jie Hu,Tianrui Li,Ghufran Ahmad Khan,Xinyan Liang,Yimiao Zhao
出处
期刊:Neurocomputing
[Elsevier BV]
日期:2021-01-09
卷期号:433: 96-107
被引量:100
标识
DOI:10.1016/j.neucom.2020.12.094
摘要
Clustering large and high-dimensional document data has got a great interest. However, current clustering algorithms lack efficient representation learning. Implementing deep learning techniques in document clustering can strengthen the learning processes. In this work, we simultaneously disentangle the problem of learned representation by preserving important information from the initial data while pushing the original samples and their augmentations together in one hand. Furthermore, we handle the cluster locality preservation issue by pushing neighboring data points together. To that end, we first introduce Contractive Autoencoders. Then we propose a deep embedding clustering framework based on contractive autoencoder (DECCA) to learn document representations. Furthermore, to grasp relevant document or word features, we append the Frobenius norm as penalty term to the conventional autoencoder framework, which helps the autoencoder to perform better. In this way, the contractive autoencoders apprehend the local manifold structure of the input data and compete with the representations learned by existing methods. Finally, we confirm the supremacy of our proposed algorithm over the state-of-the-art results on six real-world images and text datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI