异常检测
计算机科学
异常(物理)
公制(单位)
数据挖掘
人工智能
选型
集合(抽象数据类型)
系列(地层学)
秩(图论)
选择(遗传算法)
机器学习
数据集
比例(比率)
模式识别(心理学)
数学
组合数学
物理
生物
量子力学
古生物学
经济
凝聚态物理
运营管理
程序设计语言
作者
Mononito Goswami,Cristian Challú,Laurent Callot,Lenon Minorics,Andrey Kan
标识
DOI:10.48550/arxiv.2210.01078
摘要
Anomaly detection in time-series has a wide range of practical applications. While numerous anomaly detection methods have been proposed in the literature, a recent survey concluded that no single method is the most accurate across various datasets. To make matters worse, anomaly labels are scarce and rarely available in practice. The practical problem of selecting the most accurate model for a given dataset without labels has received little attention in the literature. This paper answers this question i.e. Given an unlabeled dataset and a set of candidate anomaly detectors, how can we select the most accurate model? To this end, we identify three classes of surrogate (unsupervised) metrics, namely, prediction error, model centrality, and performance on injected synthetic anomalies, and show that some metrics are highly correlated with standard supervised anomaly detection performance metrics such as the $F_1$ score, but to varying degrees. We formulate metric combination with multiple imperfect surrogate metrics as a robust rank aggregation problem. We then provide theoretical justification behind the proposed approach. Large-scale experiments on multiple real-world datasets demonstrate that our proposed unsupervised approach is as effective as selecting the most accurate model based on partially labeled data.
科研通智能强力驱动
Strongly Powered by AbleSci AI