计算机科学
预处理器
冗余(工程)
降维
维数之咒
相关性(法律)
特征选择
实施
大数据
SPARK(编程语言)
算法
并行计算
数据挖掘
人工智能
法学
操作系统
程序设计语言
政治学
作者
Sergio Ramírez‐Gallego,Iago Lastra,David Martínez‐Rego,Verónica Bolón‐Canedo,José M. Benítez,Francisco Herrera,Amparo Alonso‐Betanzos
摘要
With the advent of large-scale problems, feature selection has become a fundamental preprocessing step to reduce input dimensionality. The minimum-redundancy-maximum-relevance (mRMR) selector is considered one of the most relevant methods for dimensionality reduction due to its high accuracy. However, it is a computationally expensive technique, sharply affected by the number of features. This paper presents fast-mRMR, an extension of mRMR, which tries to overcome this computational burden. Associated with fast-mRMR, we include a package with three implementations of this algorithm in several platforms, namely, CPU for sequential execution, GPU (graphics processing units) for parallel computing, and Apache Spark for distributed computing using big data technologies.
科研通智能强力驱动
Strongly Powered by AbleSci AI