过度拟合
推车
计算机科学
数据挖掘
决策树
特征选择
人工智能
缺少数据
特征(语言学)
数据集
集合(抽象数据类型)
模式识别(心理学)
机器学习
人工神经网络
机械工程
工程类
语言学
程序设计语言
哲学
作者
Rong Tang,Xiaojun Zhang
标识
DOI:10.1109/icbda49040.2020.9101199
摘要
As the data stored in the medical database may contain missing values and redundant data, making medical data classification challenging. According to the characteristics of the medical data set containing missing values, the classification and regression (CART) algorithm is naturally thought of. However, when the CART algorithm processes a data set with too many categories, the error rate will increase rapidly and easily lead to overfitting. This paper proposes a solution for the characteristics of medical data sets and the shortcomings of CART algorithm. In order to improve the accuracy of medical data, the Boruta method was proposed to reduce the dimension. Then CART algorithm is used to classify feature subset. The data set on UCI was used in the experiment, and the results show that the accuracy of the CART algorithm is improved.
科研通智能强力驱动
Strongly Powered by AbleSci AI