可靠性(半导体)
拉什模型
计算机科学
自然语言处理
质量(理念)
随机性
人工智能
外语
语言评估
语言能力
翻译
粒度
相关性
心理学
实证研究
机器学习
有效性
语言模型
变化(天文学)
度量(数据仓库)
语言学
计算语言学
协议
统计
项目反应理论
等级间信度
标识
DOI:10.1177/02655322251406297
摘要
Assessing translation and interpreting (T&I) is essential in tertiary-level T&I education, professional certification, and foreign language testing. Recently, researchers have explored automating T&I assessment, with large language models (LLMs) emerging as a promising agent for automatic scoring. This study presents one of the first large-scale empirical investigations into the scoring reliability, severity, and validity of GPT-4o and DeepSeek-R1 in English–Chinese consecutive and simultaneous interpreting assessment. Using more than 500 pre-scored samples from the Interpreting Quality Evaluation Corpus (IQEC), the study configured eight e-raters per LLM, systematically varying three scoring parameters: reference availability (zero vs. four references), scoring granularity (segment vs. document-level scoring), and model randomness (temperature 0 vs. 1). A combination of correlation, linear mixed model, and Rasch analyses revealed that: (a) both LLMs demonstrated higher reliability than human raters; (b) DeepSeek-R1 applied significantly harsher scoring patterns than GPT-4o; (c) both LLMs achieved moderately strong correlations with human raters, with overall Spearman’s correlation coefficients ranging from .586 to .700; (d) GPT-4o exhibited higher scoring accuracy than DeepSeek-R1; and (e) LLM-based e-raters’ performance varied significantly across different scoring conditions. These results have important theoretical and practical implications, providing insights into optimizing LLM-based automatic scoring for interpreting and broader language assessment contexts.
科研通智能强力驱动
Strongly Powered by AbleSci AI