缺少数据
插补(统计学)
计算机科学
主成分分析
数据挖掘
概率逻辑
人工智能
模式识别(心理学)
机器学习
作者
Li Qu,Jianming Hu,Li Li,Yi Zhang
标识
DOI:10.1109/tits.2009.2026312
摘要
The missing data problem greatly affects traffic analysis. In this paper, we put forward a new reliable method called probabilistic principal component analysis (PPCA) to impute the missing flow volume data based on historical data mining. First, we review the current missing data-imputation method and why it may fail to yield acceptable results in many traffic flow applications. Second, we examine the statistical properties of traffic flow volume time series. We show that the fluctuations of traffic flow are Gaussian type and that principal component analysis (PCA) can be used to retrieve the features of traffic flow. Third, we discuss how to use a robust PCA to filter out the abnormal traffic flow data that disturb the imputation process. Finally, we recall the theories of PPCA/Bayesian PCA-based imputation algorithms and compare their performance with some conventional methods, including the nearest/mean historical imputation methods and the local interpolation/regression methods. The experiments prove that the PPCA method provides significantly better performance than the conventional methods, reducing the root-mean-square imputation error by at least 25%.
科研通智能强力驱动
Strongly Powered by AbleSci AI