预处理器
数据预处理
计算机科学
贝叶斯概率
采样(信号处理)
缺少数据
数据挖掘
样品(材料)
贝叶斯网络
统计
机器学习
人工智能
数学
化学
滤波器(信号处理)
色谱法
计算机视觉
作者
Runwei Li,Jacqueline MacDonald Gibson
标识
DOI:10.1021/acs.est.3c00348
摘要
The plethora of data on PFASs in environmental samples collected in response to growing concern about these chemicals could enable the training of machine-learning models for predicting exposure risks. However, differences in sampling and analysis methods across data sets must be reconciled through data preprocessing, and little information is available about how such manipulations affect the resulting models. This study evaluates how data preprocessing influences machine-learned Bayesian network models of PFOA in groundwater. We link 19 years of PFOA measurements from Minnesota, USA, to publicly available information about potential PFOA sources and factors that may influence their environmental fate. Nine different preprocessing methods were tested, and the resulting data sets were used to train models to predict the probability of PFOA ≥ 35 ppt, the 2017 Minnesota health advisory level. Different preprocessing approaches produced varying model structures with significantly different accuracies. Nonetheless, models showed similar relationships between predictor variables and PFOA exposure risks, and all models were relatively accurate, distinguishing wells at high risk from those at low risk for 82.0% to 89.0% of test data samples. There was a trade-off between data quality and model performance since a stricter data screening strategy decreased the sample size for model training.
科研通智能强力驱动
Strongly Powered by AbleSci AI