计算机科学
特征选择
对偶(语法数字)
图形
人工智能
模式识别(心理学)
选择(遗传算法)
数据挖掘
特征(语言学)
缺少数据
机器学习
理论计算机科学
语言学
文学类
哲学
艺术
作者
Zhi Qin,Hongmei Chen,Tengyu Yin,Zhong Yuan,Chuan Luo,Shi‐Jinn Horng,Tianrui Li
标识
DOI:10.1016/j.eswa.2025.126662
摘要
Graph-based feature selection methods can capture the most discriminative subset of features in multi-label data and have made great strides in recent years. However, most existing methods have three drawbacks: (1) Only binary pairwise relationship is difficult to capture latent structural information in high-dimensional data fully. (2) Using only one distance metric cannot accurately mine the similarity relationship between instances. (3) Ignore the problem that some of the non-obvious labels are missing when labels are abundant. Therefore, this paper proposes a graph diffusion with dual-distance metrics for missing multi-label feature selection named GDMMFS. GDMMFS uses both Euclidean and cosine distances to explore instance correlations and then uses diffusion techniques to transform binary relationship into quaternary relationship, thus obtaining a stable similarity graph. Moreover, GDMMFS utilizes adaptive graphs to learn latent structure information to improve the model’s immunity to interference. Finally, the ℓ 2 , 1 -norm is used as a sparse constraint to guide the weight matrix to capture the most discriminative subset of features. To further consider the possibility of missing label data, GDMMFS constructed a logic matrix to guide the numerical labels in recovering the missing information. Comparison experiments with state-of-the-art related algorithms demonstrate the superiority of the proposed algorithm, and ablation experiments demonstrate the effectiveness of graph diffusion with two-distance metrics.
科研通智能强力驱动
Strongly Powered by AbleSci AI