人工智能
计算机科学
机器学习
培训(气象学)
集成学习
样本量测定
训练集
样品(材料)
支持向量机
集合预报
数据挖掘
特征(语言学)
模式识别(心理学)
作者
Nicholas Mitsakakis,Dan Liu,Thomas Walters,Khaled El Emam
出处
期刊:Patterns
[Elsevier BV]
日期:2026-03-26
卷期号:7 (6): 101498-101498
被引量:1
标识
DOI:10.1016/j.patter.2026.101498
摘要
Health research studies often suffer from small sample sizes, and training machine learning (ML) models requires large datasets. There is a dearth of literature on determining the adequate sample size for using ML models. We developed an empirically derived sample size calculator for ensemble ML models: random forests and two gradient-boosted decision trees (light gradient boosting machine [LGBM] and extreme gradient boosting [XGBoost]). This predicts the sample size required to achieve a pre-defined level of prognostic performance with a certain probability. Prognostic performance is defined as the sample area under the ROC curve (ROC-AUC) relative to the optimal model trained on the full (population) dataset. Our calculator's accuracy was compared to three common heuristics and a statistical approach to sample size calculation. For example, the median relative error sample size prediction was 25% to achieve 85% of the optimal performance with 90% certainty for LGBM. Our model has significantly better accuracy than other methods for tree-based ensemble ML models.
科研通智能强力驱动
Strongly Powered by AbleSci AI