逻辑回归
机器学习
人工智能
预测建模
计算机科学
风险评估
算法
肺癌
统计分类
回归分析
回归
肺癌筛查
数据挖掘
医学
癌症
癌症筛查
支持向量机
作者
T. Zhang,Yexin Chen,Pei Wang,Yuting Gu,Jie Jiang,Jianxing He,Wenhua Liang
摘要
OBJECTIVE: Utilizing lung cancer risk prediction models at the screening stage can enhance the accuracy of identifying high-risk individuals eligible for lung cancer screening. However, there is a relative lack of research on such prediction models in China, particularly regarding machine learning algorithms. METHODS: A stratified random sampling method was employed to randomly divide the dataset into a training set (70%) and a validation set (30%). Key variables were screened using LASSO regression. Then logistic regression and XGBoost algorithm were utilized to construct a lung cancer risk prediction model in the training set and validate it in the validation set, respectively. RESULTS: A lung cancer risk prediction model was constructed using 11,708 participants enrolled in a prospective cohort, the Guangzhou Lung-Care Project Program. In the constructed lung cancer risk prediction models, the AUC of the logistic regression model in the validation set was 0.647 (95% CI: 0.574-0.720); in contrast, the AUC of the XGBoost model based on the machine learning algorithm in the validation set was 0.658 (95% CI: 0.589-0.727), demonstrating slightly better discriminative ability compared to the logistic regression model. In addition, this study found the important effect of childhood exposure to cooking fuels on the risk of lung cancer, which has been rarely considered in previous research. CONCLUSION: The lung cancer risk prediction model constructed based on the XGBoost algorithm is better than the logistic regression algorithm in terms of prediction accuracy and robustness, aiding in the risk assessment of individuals undergoing screening.
科研通智能强力驱动
Strongly Powered by AbleSci AI