计算机科学
机器学习
快照(计算机存储)
人工智能
差异(会计)
人口
水准点(测量)
训练集
一般化
主动学习(机器学习)
数据挖掘
置信区间
采样(信号处理)
航程(航空)
低信心
过程(计算)
分类器(UML)
注释
合成数据
稳健性(进化)
总体方差
数据建模
统计
集成学习
监督学习
作者
Tianjiao Wan,Zijian Gao,Xudong Gong,Dawei Feng,Xingxing Zhang,Bo Ding,Yijie Wang,Huaimin Wang,Kele Xu
标识
DOI:10.1109/tcsvt.2025.3622313
摘要
Active Learning (AL) aims to reduce data annotation costs by selecting the most informative samples from an unlabeled data pool. Traditional AL methods often rely on a single snapshot to identify uncertain or representative samples, often overlooking the poor generalization of a single model. Recent AL studies have attempted to address this issue by tracking a broader range of training dynamics for data selection, typically using averaging or accumulating manner. However, both our theoretical and experimental analyses reveal that these methods obscure the variability inherent in the training process, potentially prioritizing hard-to-learn samples that result in poor generalization. In this paper, we propose a novel AL method termed as Dynamic Confidence Variance (DCoV), that seamlessly integrates variability with the training dynamic to effectively identify a well-generalized Coreset. DCoV leverages the variance of the model’s prediction confidence throughout the training process for active sampling and model training. Our theoretical analysis demonstrates that DCoV provides a lower bound on the population risk of the model learned from selected labeled subset, spanning the entire training process. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art AL methods on various balanced and imbalanced benchmark datasets across various modalities.
科研通智能强力驱动
Strongly Powered by AbleSci AI