催化作用
化学
计算机科学
组合化学
人工智能
程序设计语言
软件工程
自然语言处理
自然语言
领域(数学分析)
作者
Honghao Chen,Hongxuan Liu,Yishen Tew,Xiaotian Ren,Xiaojin Tang,Xiaonan Wang
出处
期刊:ACS Catalysis
[American Chemical Society]
日期:2025-10-20
卷期号:15 (21): 18244-18254
被引量:17
标识
DOI:10.1021/acscatal.5c06431
摘要
Decades of catalysis knowledge remain locked in unstructured prose, hindering data-driven discovery. Existing text-mining tools struggle to establish the synthesis–structure–performance relationships critical for catalyst knowledge discovery as they rarely connect synthesis protocols in one section with the resulting material properties and performance outcomes reported elsewhere. Here, we present CATDA (Corpus-aware Automated Text-to-Graph Catalyst Discovery Agent), a long-context large language model-driven agentic framework that reads full documents and distills them into actionable, provenance-tracked knowledge graphs linking material properties, multistep synthesis, conditions, and testing outcomes. Applied at corpus scale, CATDA extracts data with near-human fidelity (F1 = 0.983) and a 12-fold speedup over manual curation. This structured knowledge is made accessible through two synergistic applications: a DatasetAgent for exporting machine-learning-ready tables and a CatAgent providing a conversational, citation-linked interface for interactive discovery. The high-quality data set enabled the training of a predictive model for ethylbenzene conversion while simultaneously exposing systemic challenges such as feature sparsity and protocol heterogeneity in the source literature. By transforming the literature into a queryable and computable resource, CATDA offers a scalable route to accelerate large-scale data analysis, quantitative modeling, and a rational catalyst design paradigm.
科研通智能强力驱动
Strongly Powered by AbleSci AI