水准点(测量)
计算机科学
人工智能
科学推理
感知
数据科学
科学文献
语言模型
自然语言处理
科学建模
视觉语言
可视化
编码器
机器学习
认知科学
实验数据
科学发现
计算模型
科学写作
作者
Hanzheng Li,Xi Fang,Yixuan Li,Chaozheng Huang,Yì Wáng,Xi Wang,Hongzhe Bai,Bojun Hao,Shenyu Lin,Huiqi Liang,Linfeng Zhang,Guolin Ke
标识
DOI:10.1021/acs.jcim.6c00286
摘要
The integration of multimodal large language models (MLLMs) into chemistry promises to revolutionize scientific discovery, yet their ability to comprehend the dense, graphical language of reactions within authentic literature remains underexplored. Here, we introduce RxnBench, a multitiered benchmark designed to rigorously evaluate MLLMs on chemical reaction understanding from scientific PDFs. RxnBench comprises two tasks: Single-figure QA (SF-QA), which tests fine-grained visual perception and mechanistic reasoning using 1525 questions derived from 305 curated reaction schemes, and Full-Document QA (FD-QA), which challenges models to synthesize information from 108 articles, requiring cross-modal integration of text, schemes, and tables. Our evaluation of MLLMs reveals a critical capability gap: while models excel at extracting explicit text, they struggle with deep chemical logic and precise structural recognition. Notably, models with inference-time reasoning significantly outperform standard architectures, yet none achieve 50% accuracy on FD-QA. These findings underscore the urgent need for domain-specific visual encoders and stronger reasoning engines to advance autonomous AI chemists.
科研通智能强力驱动
Strongly Powered by AbleSci AI