随机森林
计算机科学
特征工程
机器学习
人工智能
学习分析
梯度升压
支持向量机
Boosting(机器学习)
任务(项目管理)
二元分类
集成学习
预测分析
特征(语言学)
分析
预测建模
原始数据
数据分析
数据科学
集合预报
大数据
二进制数
人工神经网络
深度学习
特征学习
动力学(音乐)
特征向量
任务分析
摘要
ABSTRACT This study proposes a theory‐informed learning analytics approach for predicting at‐risk students using large‐scale behavioural, demographic and academic data. Building on engagement theory and self‐regulated learning, we engineer pedagogically grounded behavioural indicators that move beyond raw click counts. These indicators include multidimensional engagement measures, temporal regularity and an antiprocrastination score derived from assessment submission patterns. Using the Open University Learning Analytics Dataset (OULAD), comprising 32,593 students across 22 courses, we reformulate the prediction task as a binary classification problem (Pass vs. At‐Risk) and compare three machine learning algorithms: Support Vector Machine (SVM), Random Forest (RF) and Extreme Gradient Boosting (XGBoost). Models are evaluated at four quarter‐based checkpoints over the semester to investigate temporal dynamics and opportunities for timely intervention. Results show that XGBoost consistently outperforms RF and SVM in accuracy, recall, precision and ROC AUC, while behavioural features overwhelmingly dominate demographic and academic variables in predictive importance. Temporal analysis reveals that model performance improves substantially from the first to the third quarter, with mid‐semester predictions offering the best trade‐off between accuracy and time remaining for effective support. The findings demonstrate the value of theory‐driven feature engineering and temporally sensitive evaluation in designing early‐warning systems that are both accurate and pedagogically actionable.
科研通智能强力驱动
Strongly Powered by AbleSci AI