Comparing programming languages for data analytics: Accuracy of estimation in Python and R

Python(编程语言) 计算机科学 分析 程序设计语言 数据挖掘
作者
Chelsey Hill,Lanqing Du,Marina Johnson,B. D. McCullough
出处
期刊:Wiley Interdisciplinary Reviews-Data Mining and Knowledge Discovery [Wiley]
卷期号:14 (3) 被引量:8
标识
DOI:10.1002/widm.1531
摘要

Abstract Several open‐source programming languages, particularly R and Python, are utilized in industry and academia for statistical data analysis, data mining, and machine learning. While most commercial software programs and programming languages provide a single way to deliver a statistical procedure, open‐source programming languages have multiple libraries and packages offering many ways to complete the same analysis, often with varying results. Applying the same statistical method across these different libraries and packages can lead to entirely different solutions due to the differences in their implementations. Therefore, reliability and accuracy should be essential considerations when making library and package usage decisions while conducting statistical analysis using open source programming languages. Instead, most users take this for granted, assuming that their chosen libraries and packages produce accurate results for their statistical analysis. To this extent, this study assesses the estimation accuracy and reliability of Python and R's various libraries and packages by evaluating the univariate summary statistics, analysis of variance (ANOVA), and linear regression procedures using benchmarking data from the National Institutes of Standards and Technology (NIST). Further, experimental results are presented comparing machine learning methods for classification and regression. The libraries and packages assessed in this study include the stats package in R and Pandas, Statistics, NumPy, statsmodels, SciPy, statsmodels, scikit‐learn, and pingouin in Python. The results show that the stats package in R and statsmodels library in Python are reliable for univariate summary statistics. In contrast, Python's scikit‐learn library produces the most accurate results and is recommended for ANOVA. Among the libraries and packages assessed for linear regression, the results demonstrated that the stats package in R is more reliable, accurate, and flexible; thus, it is recommended for linear regression analysis. Further, we present results and recommendations for machine learning using R and Python. This article is categorized under: Algorithmic Development > Statistics Application Areas > Data Mining Software Tools
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
1秒前
此时此刻完成签到 ,获得积分10
5秒前
5秒前
hugeyoung发布了新的文献求助10
5秒前
5秒前
风趣的翼发布了新的文献求助30
6秒前
6秒前
7秒前
7秒前
9秒前
10秒前
Battery-Li发布了新的文献求助10
11秒前
11秒前
酷波er的应助被lb采纳,获得10
12秒前
帝轩泽发布了新的文献求助30
12秒前
ding的应助被从未伤害你采纳,获得10
13秒前
13秒前
spiritMaxs发布了新的文献求助20
14秒前
轻松悒完成签到 ,获得积分10
14秒前
00发布了新的文献求助10
14秒前
芒果发布了新的文献求助10
15秒前
奚瑞发布了新的文献求助10
15秒前
cdercder的应助被DND采纳,获得10
16秒前
耳机杀手发布了新的文献求助10
16秒前
搜集达人的应助被科研通管家采纳,获得10
16秒前
搜集达人的应助被科研通管家采纳,获得10
16秒前
星辰大海的应助被科研通管家采纳,获得10
16秒前
科研通AI2S的应助被科研通管家采纳,获得10
17秒前
酷波er的应助被科研通管家采纳,获得10
17秒前
汉堡包的应助被科研通管家采纳,获得10
17秒前
清新的应助被科研通管家采纳,获得10
17秒前
天天快乐的应助被科研通管家采纳,获得10
17秒前
深情安青的应助被科研通管家采纳,获得10
17秒前
aajhajkahna的应助被科研通管家采纳,获得10
17秒前
香蕉觅云的应助被科研通管家采纳,获得10
17秒前
aajhajkahna的应助被科研通管家采纳,获得10
18秒前
烟花的应助被科研通管家采纳,获得10
18秒前
顾矜的应助被科研通管家采纳,获得10
18秒前
共享精神的应助被科研通管家采纳,获得10
18秒前
清新的应助被科研通管家采纳,获得10
18秒前
高分求助中
(应助此贴封号)通过应助OA文献获取积分 10000
Rosenblum, Global Change Biology 800
The Art of Interactive Teaching 600
Computational Chemical Reaction Engineering: Modeling, Simulation, and Design with MATLAB 600
Organizational Behavior 510
Management and the Arts 510
CLSI C56QG Examples of Hemolyzed, Icteric, and Lipemic/Turbid Samples Quick Guide 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 计算机科学 工程类 纳米技术 内科学 物理 有机化学 化学工程 生物化学 复合材料 光电子学 细胞生物学 心理学 量子力学 催化作用 物理化学 电极
热门帖子
关注 科研通微信公众号,转发送积分 7800868
求助须知:如何正确求助?哪些是违规求助? 9335531
关于积分的说明 20474429
捐赠科研通 7392513
什么是DOI,文献DOI怎么找? 3326463
关于科研通互助平台的介绍 2473394
邀请新用户注册赠送积分活动 2344280