SPARK(编程语言)
支持向量机
大数据
计算机科学
相关向量机
维数(图论)
结构化支持向量机
机器学习
分拆(数论)
数据挖掘
人工智能
样品(材料)
数学
化学
色谱法
组合数学
纯数学
程序设计语言
作者
Siyang Lu,Yihong Chen,Zhu Xiao-lin,Ziyi Wang,Yangjun Ou,Yuhang Xie
标识
DOI:10.1145/3494885.3494891
摘要
The traditional support vector machines perform well in classification and prediction on small and medium-sized data sets, but there are some problems such as the low training efficiency and the low accuracy in large sample number, high dimension and large-scale data sets. Meanwhile, with the rise of distributed computing platforms such as the Spark suitable for big data analyses, more and more scholars at home and abroad turn their research direction to the distributed machine learning algorithms Therefore, in order to carry out the research on support vector machine for big data analyses, this paper explores the related researches and current situations of support vector machine, including: in-depth analysis of the algorithm principle of support vector machines, systematical investigation of the improved methods of support vector machines for the big data analyses, and distributed support vector machines under the Spark platform. Then, combined with the parallelization mechanism of the Spark, the some future research directions of support vector machine are investigated: for optimizing the accuracy of training results, some special matrix calculation skills should be added; In term of the research on SVM under the Spark platform, some better optimization methods from the perspective of dimension and partition can be found.
科研通智能强力驱动
Strongly Powered by AbleSci AI