注释
预处理器
计算机科学
数据集成
计算生物学
情报检索
数据挖掘
人工智能
生物
标识
DOI:10.1186/s13059-025-03639-x
摘要
Abstract Single-cell RNA sequencing has revolutionized cellular heterogeneity research, but analyzing the abundance of unannotated public datasets remains challenging. We present scExtract, a framework leveraging large language models to automate scRNA-seq data analysis from preprocessing to annotation and integration. scExtract extracts information from research articles to guide data processing, outperforming existing reference transfer methods in benchmarks. We introduce scanorama-prior and cellhint-prior, which incorporate prior annotation information for improved batch correction while preserving biological diversities. We demonstrate scExtract’s utility by integrating 14 datasets to create a comprehensive human skin atlas of 440,000 cells.
科研通智能强力驱动
Strongly Powered by AbleSci AI