| 标题 |
Evaluating the reasoning capabilities of large language models in Chinese-language contexts
|
| 网址 | |
| DOI |
暂未提供,该求助的时间将会延长,查看原因?
|
| 其它 |
Abstract With the rapid iteration of AI technologies, reasoning capabilities have become a core indicator for measuring the intelligence level of large language models (LLMs) and a focus of research in both academia and industry. This report aims to establish a systematic, objective, and comprehensive evaluation framework to assess AI reasoning capabilities. We compared 36 LLMs on various text-based reasoning tasks in Chinese-language contexts and found that GPT-o3 achieved the highest score in the basic logical reasoning evaluation, while Gemini 2.5 Flash led in contextual reasoning evaluation. In terms of overall ranking, Doubao 1.5 Pro (Thinking) secured the top position, closely followed by OpenAI’s recently released GPT-5 (Auto). Several Chinese-developed LLMs—including Doubao 1.5 Pro, Qwen 3 (Thinking), |
| 求助人 | |
| 下载 | 该求助完结已超 24 小时,文件已从服务器自动删除,无法下载。 |
PDF的下载单位、IP信息已删除
(2025-6-4)