Are large language models superhuman chemists?

高分子科学 化学
作者
Adrian Mirza,Nawaf Alampara,Sreekanth Kunchapu,Benedict Emoekabu,Aswanth Krishnan,Mara Wilhelmi,Macjonathan Okereke,J. Eberhardt,Amir Mohammad Elahi,Maximilian Greiner,Caroline T. Holick,Tanya Gupta,Mehrdad Asgari,Christina Glaubitz,Lea C. Klepsch,Yannik Köster,Jakob Meyer,Santiago Miret,Tim Hoffmann,Fabian Alexander Kreth
出处
期刊:Cornell University - arXiv [Cornell University]
被引量:13
标识
DOI:10.48550/arxiv.2404.01475
摘要

Large language models (LLMs) have gained widespread interest due to their ability to process human language and perform tasks on which they have not been explicitly trained. This is relevant for the chemical sciences, which face the problem of small and diverse datasets that are frequently in the form of text. LLMs have shown promise in addressing these issues and are increasingly being harnessed to predict chemical properties, optimize reactions, and even design and conduct experiments autonomously. However, we still have only a very limited systematic understanding of the chemical reasoning capabilities of LLMs, which would be required to improve models and mitigate potential harms. Here, we introduce "ChemBench," an automated framework designed to rigorously evaluate the chemical knowledge and reasoning abilities of state-of-the-art LLMs against the expertise of human chemists. We curated more than 7,000 question-answer pairs for a wide array of subfields of the chemical sciences, evaluated leading open and closed-source LLMs, and found that the best models outperformed the best human chemists in our study on average. The models, however, struggle with some chemical reasoning tasks that are easy for human experts and provide overconfident, misleading predictions, such as about chemicals' safety profiles. These findings underscore the dual reality that, although LLMs demonstrate remarkable proficiency in chemical tasks, further research is critical to enhancing their safety and utility in chemical sciences. Our findings also indicate a need for adaptations to chemistry curricula and highlight the importance of continuing to develop evaluation frameworks to improve safe and useful LLMs.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
搜集达人应助风吹采纳,获得10
2秒前
sunsuan发布了新的文献求助10
3秒前
whisper应助咿呀采纳,获得10
4秒前
赘婿应助冷酷的依霜采纳,获得10
4秒前
不吃辣椒发布了新的文献求助10
5秒前
5秒前
haitun发布了新的文献求助10
6秒前
6秒前
简单面包完成签到,获得积分10
6秒前
坚强的烤鸡完成签到,获得积分20
7秒前
酷波er应助科研文献采纳,获得10
7秒前
7秒前
7秒前
黑暗暴龙神完成签到,获得积分10
7秒前
qingqingdandan完成签到 ,获得积分10
8秒前
林夕关注了科研通微信公众号
8秒前
xi完成签到,获得积分10
8秒前
9秒前
优美的高山完成签到,获得积分10
9秒前
青鸟发布了新的文献求助10
10秒前
乐乐应助tzzwa采纳,获得20
10秒前
勇yi发布了新的文献求助10
10秒前
wanci应助陈栩采纳,获得10
10秒前
霸气冰露完成签到,获得积分10
11秒前
隐形曼青应助陈一会采纳,获得10
12秒前
扎心发布了新的文献求助10
12秒前
日月归尘发布了新的文献求助10
12秒前
汉堡包应助安静的幼旋采纳,获得10
13秒前
xi发布了新的文献求助10
13秒前
orixero应助孙伟健采纳,获得10
14秒前
14秒前
酷酷雁枫发布了新的文献求助10
14秒前
最最完成签到,获得积分10
14秒前
dd完成签到,获得积分20
14秒前
ALLEN发布了新的文献求助10
15秒前
orixero应助怕黑的惜霜采纳,获得10
15秒前
青鸟完成签到,获得积分20
17秒前
17秒前
kurii发布了新的文献求助10
18秒前
18秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Römisch-Germanische Forschungen 1000
China Pluperfect I: Epistemology of Past and Outside in Chinese Art 520
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The fast track to determining transfer functions of linear circuits: The student guide 500
The Analytical and Numerical Solution of Electric and Magnetic Fields 500
Green Fire Retardants for Polymeric Materials 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7617475
求助须知:如何正确求助?哪些是违规求助? 9192742
关于积分的说明 19701410
捐赠科研通 7189743
什么是DOI,文献DOI怎么找? 3272020
关于科研通互助平台的介绍 2434795
邀请新用户注册赠送积分活动 2267123