概念漂移
鉴别器
计算机科学
流式数据
星团(航天器)
人工智能
数据挖掘
无监督学习
集合(抽象数据类型)
数据集
钥匙(锁)
模式识别(心理学)
机器学习
数据流挖掘
探测器
程序设计语言
电信
计算机安全
作者
Mingjie Zhao,Yiqun Zhang,Yuzhu Ji,Yang Lu
标识
DOI:10.1007/978-981-99-8435-0_3
摘要
Streaming data processing has attracted much more attention and become a key research area in the fields of machine learning and data mining. Since the distribution of real data may evolve (called concept drift) with time due to many unforeseen factors and real data is usually with imbalanced cluster/class distributions during streaming data processing, drifts occurred in distributions with fewer data objects are easily masked by the larger distributions. This paper, therefore, proposes an unsupervised drift detection approach called Multi-Imbalanced Cluster Discriminator (MICD) to address the more challenging imbalance problem of unlabeled data. It first partitions data into compact clusters, and then learns a discriminator for each cluster to detect drift. It turns out that MICD can detect drift occurrence, locate where the drift occurs, and quantify the extent of the drift. MICD is efficient, interpretable, and has easy-to-set parameters. Extensive experiments on synthetic and real datasets illustrate the superiority of MICD.
科研通智能强力驱动
Strongly Powered by AbleSci AI