超参数
计算机科学
人工智能
机器学习
特征选择
超参数优化
分类器(UML)
支持向量机
鉴定(生物学)
数据挖掘
多类分类
集成学习
集合预报
选型
统计分类
特征(语言学)
可扩展性
模式识别(心理学)
选择(遗传算法)
精确性和召回率
随机森林
统计模型
一般化
特征工程
编码(社会科学)
线性分类器
特征提取
交叉验证
线性判别分析
监督学习
作者
Mutiana Pratiwi,Sarjon Defit,Muhammad Tajuddin
标识
DOI:10.1109/icodsa67155.2025.11157136
摘要
This research proposes an optimized framework named the TCO (Text Classification Optimization) model, which synergizes hyperparameter tuning and feature selection to enhance multiclass text classification tasks, specifically for automated competency identification in recruitment systems. The framework employs Term Frequency–Inverse Document Frequency (TF-IDF) for robust feature extraction, Chi-Square statistical testing for relevant feature selection, and Optuna for efficient hyperparameter optimization. This integrated approach addresses the challenges of high-dimensionality and model generalization in textual datasets. To further improve classification performance, the TCO model implements an ensemble stacking strategy that combines the strengths of multiple base classifiers, including Support Vector Classifier (SVC), Naïve Bayes, Random Forest, Gradient Boosting, and XGBoost. Experimental results show that the combination of optimized hyperparameters and selected features significantly boosts model performance in comparison to baseline models without such enhancements. The TCO model achieved an overall classification accuracy of 74%, with notable precision (86%) in identifying the "Problem Solving" competency class. Balanced recall and F1-score metrics across classes indicate consistent performance and reduced bias. This research contributes meaningfully to the development of fair, interpretable, and scalable automated recruitment systems. By leveraging data-driven methods and ensemble learning, the TCO model offers a reliable solution for organizations seeking to streamline talent acquisition processes while ensuring classification accuracy and competency alignment.
科研通智能强力驱动
Strongly Powered by AbleSci AI