计算机科学
情报检索
管道(软件)
工作流程
预处理器
数据挖掘
元数据
样品(材料)
关系数据库
数据库
解析
数据检索
原始数据
重新使用
数据集成
源代码
编码(集合论)
索引(排版)
语义网
匹配(统计)
Web服务
数据整理
比例(比率)
本体论
注释
人工智能
数据预处理
数据存取
作者
Yingying Zhao,Quanyou Cai,Dongzhu Chen,Jiekai Chen
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2026-06-10
标识
DOI:10.64898/2026.06.06.730646
摘要
Datasets in the Gene Expression Omnibus (GEO) remain difficult to reuse at scale because sample annotations are heterogeneous and raw sequencing data require assay-specific preprocessing. We present GEOAgent, an AI-driven autonomous framework designed for intelligent dataset retrieval and standardized preprocessing by coupling autonomous semantic governance with an automated Nextflow pipeline named bioStream. Metadata from 181,760 sequencing series and 84,756 associated PubMed records were organized in a relational database and semantic index to support natural-language dataset retrieval. The framework automatically determines assay modalities, resolves experimental design pairings, and standardizes sample naming to minimize manual curation overhead. Based on these parsed attributes, the framework generates deployment-ready manifests to automatically execute containerized workflows across bulk and single-cell omics modalities. In expert-curated benchmarks, the workflow achieved 96% retrieval precision alongside 100% accuracy in assay classification and sample relationship resolution. The web platform is publicly accessible, while the source code and associated databases are openly available via GitHub and Zenodo.
科研通智能强力驱动
Strongly Powered by AbleSci AI