元数据
生物
蓝图
基因组学
计算生物学
转录组
DNA测序
计算机科学
数据整理
序列(生物学)
数据科学
功能基因组学
万维网
信息存储库
数据存取
从头转录组组装
计算基因组学
RNA序列
情报检索
基因预测
钥匙(锁)
基因
生物学数据
传播
公共通道
信息抽取
注释
数据提取
生物信息学
顺序装配
作者
Nicholas D. Youngblut,Christopher Carpenter,Arshia Nayebnazar,Abhinav Adduri,Rohan Shah,Chiara Ricci-Tam,Jaanak Prashar,Rajesh Ilango,Noam Teyssier,Silvana Konermann,Patrick D. Hsu,Alexander Dobin,Dave P. Burke,Hani Goodarzi,Yusuf Roohani
出处
期刊:Cell
[Cell Press]
日期:2026-09-01
卷期号:189 (19): 5932-5944.e6
标识
DOI:10.1016/j.cell.2026.08.025
摘要
Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.
科研通智能强力驱动
Strongly Powered by AbleSci AI