德国的
水准点(测量)
基础(证据)
边疆
医学教育
医学院
医学
计算机科学
价值(数学)
公共卫生
政治学
心理学
执照
公共关系
残余物
委派
教育测量
家庭医学
梅德林
作者
Lasse Cirkel,Johannes Knitza,Volker Schillings,A. Oksche,Jan C. Becker,Sebastian Kuhn
标识
DOI:10.1038/s41746-026-03082-7
摘要
We evaluated proprietary and open-weight foundation models on 24 German medical licensing examinations (2019-2024), including 7485 items and response data from 119,878 sittings. For fair comparison, the eight vision-capable models were evaluated on the full benchmark and all thirteen on a shared text-only subset. On the full benchmark, Gemini 3.1 Pro achieved the highest overall accuracy, reaching 99.31% on the first (M1) and 98.37% on the second (M2) examination. On the shared text-only subset, proprietary frontier models performed at near-ceiling levels, several open-weight models (including GLM-5 and DeepSeek V3.2-Thinking) were highly competitive, and even compact ones exceeded mean student performance. Image-present items were more difficult for both students and models, but the associated decline was disproportionately larger for models than for students. Human- and model-defined difficulty subsets showed limited overlap, and model-hard subsets revealed residual differences among top systems. These findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure. They carry implications for high-stakes assessment and AI-assisted medical education, notably multimodal assessment, human-aligned educational tools, and privacy-preserving local deployment.
科研通智能强力驱动
Strongly Powered by AbleSci AI