计算机科学
聚类分析
数据挖掘
约束(计算机辅助设计)
集合(抽象数据类型)
数据集
领域(数学分析)
人工智能
数学
几何学
数学分析
程序设计语言
作者
Lecheng Zheng,Yada Zhu,Jingrui He
出处
期刊:Society for Industrial and Applied Mathematics eBooks
[Society for Industrial and Applied Mathematics]
日期:2023-01-01
卷期号:: 856-864
被引量:9
标识
DOI:10.1137/1.9781611977653.ch96
摘要
In the era of big data, we are often facing the challenge of data heterogeneity and the lack of label information simultaneously. In the financial domain (e.g., fraud detection), the heterogeneous data may include not only numerical data (e.g., total debt and yearly income), but also text and images (e.g., financial statement and invoice images). At the same time, the label information (e.g., fraud transactions) may be missing for building predictive models. To address these challenges, many state-of-the-art multi-view clustering methods have been proposed and achieved outstanding performance. However, these methods typically do not take into consideration the fairness aspect and are likely to generate biased results using sensitive information such as race and gender. Therefore, in this paper, we propose a fairness-aware multi-view clustering method named Fair-MVC. It incorporates the group fairness constraint into the soft membership assignment for each cluster to ensure that the fraction of different groups in each cluster is approximately identical to the entire data set. Meanwhile, we adopt the idea of both contrastive learning and non-contrastive learning and propose novel regularizers to handle heterogeneous data in complex scenarios with missing data or noisy features. Experimental results on real-world data sets demonstrate the effectiveness and efficiency of the proposed framework. We also derive insights regarding the relative performance of the proposed regularizers in various scenarios.KeywordsMulti-view LearningContrastive LearningClustering
科研通智能强力驱动
Strongly Powered by AbleSci AI