过度拟合
预处理器
计算机科学
蓝图
模型验证
机器学习
预测建模
人工智能
数据挖掘
数据验证
数据预处理
工作(物理)
背景(考古学)
训练集
数据科学
交叉验证
计算模型
作者
Eneko López,Giulia Gorla,Jaione Etxebarria‐Elezgarai,José Manuel Amigo,Andreas Seifert
标识
DOI:10.1016/j.aca.2025.344838
摘要
Overfitting remains one of the most pervasive and deceptive pitfalls in predictive modeling. It leads to models that perform exceptionally well on training data but cannot be transferred nor generalized to real-world scenarios. Although overfitting is usually attributed to excessive model complexity, it is often the result of inadequate validation strategies, faulty data preprocessing and biased model selection, problems that can inflate apparent accuracy and compromise predictive reliability. In this second part of our series, we examine the most common yet overlooked practices that contribute to overfitting, ranging from data leakage in preprocessing to the pressures of scientific publishing that encourage result-driven overoptimization. By identifying these pitfalls and providing practical guidelines for performing robust validation protocols, this work serves as a blueprint for researchers to ensure their models are not only high-performing but also trustworthy, reproducible, and generalizable. • Comprehensive tutorial on external validation • Focus on overfitting as one of the most pervasive and deceptive pitfalls • An eight-step checklist for reliable chemometric modeling • Overfitting as a result of a chain of avoidable missteps
科研通智能强力驱动
Strongly Powered by AbleSci AI