计算机科学
特征选择
数据挖掘
比例(比率)
选择(遗传算法)
依赖关系(UML)
过程(计算)
任务(项目管理)
特征(语言学)
人工智能
模式识别(心理学)
管理
哲学
语言学
操作系统
物理
量子力学
经济
作者
Tian Yang,Yanfang Deng,Bin Yu,Yuhua Qian,Jianhua Dai
标识
DOI:10.1109/tkde.2022.3181208
摘要
Processing large-scale data sets with limited labels has always been a difficult task in data mining. Facing this difficulty, two local feature selection algorithms, LARD and LRSD, have been proposed based on dependency degree, which can process partially labeled data sets and greatly improve the computational efficiency. However, it is very difficult for these algorithms to calculate large-scale data with millions of samples on a typical personal computer. Although the related family method is a more efficient approach than dependency degree, it cannot be used for partially labeled large-scale data. As a result, a local feature selection method based on related family is proposed to accelerate data processing in the paper. Experiments show that the proposed algorithm can run 405 times faster than LARD on partially labeled data sets and maintain high classification accuracy. In addition, this new algorithm can effectively process partially labeled large-scale data sets with 5,000,000 samples or 20,000 features on a typical personal computer.
科研通智能强力驱动
Strongly Powered by AbleSci AI