元数据
计算机科学
元数据仓储
数据挖掘
数据库
数据提取
集合(抽象数据类型)
数据集
重新使用
数据文件
数据元素
样品(材料)
情报检索
可靠性(半导体)
鉴定(生物学)
信息存储库
数据采集
可重用性
参考数据
实验数据
数据验证
限制
数据质量
数据管理
关系数据库
数据映射
数据集成
数据类型
试验数据
轨道轨道
作者
Marie Andken,Clarissa Zheng,Zhi Sun,Eric W. Deutsch
标识
DOI:10.1021/acs.jproteome.5c01045
摘要
The reusability of proteomics data sets depends on the ability to obtain accurate metadata to guide reprocessing pipelines. However, many data sets deposited in public data repositories lack sufficient and reliable annotation, limiting large-scale reanalyses. To address this challenge, we developed RunAssessor, a tool that systematically extracts and summarizes information directly from mass spectrometry data files prior to peptide identification analysis. RunAssessor extracts and summarizes sample preparation and instrument acquisition parameters directly from the data where possible. Using one complete data set and test files from 18 other data sets as examples, we demonstrate RunAssessor's ability to extract instrument models, isobaric labels, phosphoenrichment, precursor and fragment ion tolerances, along with the dynamic exclusion time used by the instrument. These extracted metadata are stored in a comprehensive output file, and summarized in a standard Sample and Data Relationship Format (SDRF) file, thereby reducing the burden of manual curation and improving the reliability of proteomics data set metadata, facilitating the reuse of public data.
科研通智能强力驱动
Strongly Powered by AbleSci AI