一致性(知识库)
工作流程
计算机科学
软件
RNA序列
标杆管理
聚类分析
透明度(行为)
数据挖掘
人口
原始数据
降维
基因
数据库
生物
基因表达
人工智能
遗传学
程序设计语言
社会学
业务
人口学
转录组
营销
计算机安全
作者
Joseph M. Rich,Lambda Moses,Pétur Helgi Einarsson,Kayla Jackson,Laura Luebbert,A. Sina Booeshaghi,Sindri Emmanúel Antonsson,Delaney K. Sullivan,Nicolas Bray,Páll Melsted,Lior Pachter
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2024-04-05
被引量:35
标识
DOI:10.1101/2024.04.04.588111
摘要
Standard single-cell RNA-sequencing analysis (scRNA-seq) workflows consist of converting raw read data into cell-gene count matrices through sequence alignment, followed by analyses including filtering, highly variable gene selection, dimensionality reduction, clustering, and differential expression analysis. Seurat and Scanpy are the most widely-used packages implementing such workflows, and are generally thought to implement individual steps similarly. We investigate in detail the algorithms and methods underlying Seurat and Scanpy and find that there are, in fact, considerable differences in the outputs of Seurat and Scanpy. The extent of differences between the programs is approximately equivalent to the variability that would be introduced in benchmarking scRNA-seq datasets by sequencing less than 5% of the reads or analyzing less than 20% of the cell population. Additionally, distinct versions of Seurat and Scanpy can produce very different results, especially during parts of differential expression analysis. Our analysis highlights the need for users of scRNA-seq to carefully assess the tools on which they rely, and the importance of developers of scientific software to prioritize transparency, consistency, and reproducibility for their tools.
科研通智能强力驱动
Strongly Powered by AbleSci AI