计算机科学
机器学习
人工智能
可扩展性
过程(计算)
一般化
领域(数学分析)
数据科学
小数据
领域知识
学习迁移
稀缺
开放式研究
大概是正确的学习
大数据
数据建模
软件
训练集
标识
DOI:10.1186/s40537-025-01346-9
摘要
Abstract The abundance of large datasets has driven machine learning (ML) model performance and scalability breakthroughs. However, many domains and practical applications must contend with the limitations imposed by small and very small datasets. This survey thoroughly examines state-of-the-art methodologies and challenges in ML approaches tailored for scenarios where data scarcity is a fundamental constraint. We begin by outlining the theoretical foundations that govern learning from small data. Then, we discuss recent advancements in data-related frameworks (i.e., training and evaluation methods, etc.) and algorithmic architectures (meta and transfer learning). We also explore the trade-offs and related issues inherent in designing models for small data, such as overfitting, generalization error, and the bias-variance dilemma, as well as identify minimal interventions that can overcome such issues. Further, this survey covers the role of synthetic data generation and simulation-based approaches to enlarge data availability while critically assessing the implications of these techniques on model performance. Finally, in synthesizing open literature, we shed light on emerging trends/research directions that aim to overcome challenges arising from limited data, such as incorporating domain knowledge and causal principles to guide the learning process and integrating symbolic reasoning with statistical learning.
科研通智能强力驱动
Strongly Powered by AbleSci AI