心理学
罗夏测验
可靠性(半导体)
卡帕
协议(科学)
一致性(知识库)
内容有效性
心理测量学
等级间信度
临床心理学
测试有效性
自然语言处理
社会心理学
有效性
标准效度
应用心理学
增量有效性
心理测试
认知心理学
内部一致性
人工智能
作者
Ruam Pedro Francisco de Assis Pimentel,Gregory J. Meyer
出处
期刊:Assessment
[SAGE Publishing]
日期:2026-06-18
卷期号:: 10731911261455137-10731911261455137
标识
DOI:10.1177/10731911261455137
摘要
Large language models (LLMs) are increasingly used to support psychological assessment, but standards for evaluating their scoring accuracy remain limited. This article introduces a clear, reproducible validation framework to evaluate LLM-based scoring systems. The framework separates pre-validation steps (e.g., balancing base rates, refining prompts, and comparing models) from a standardized validation phase focused on reliability and validity benchmarks. We demonstrate its application with a case study of Morbid Content (MOR) scoring in the Rorschach task, using a two-agent LLM workflow. In an independent dataset ( n = 84; 2,176 responses) with natural MOR base rates, the final LLM coder showed good response level agreement ( kappa = .72–.74) and excellent protocol level agreement ( ICC = 0.94–0.95) with assessors, near-perfect consistency with itself (ICC = 0.97–0.99), and replicated external validity ( r = .59–.71) that matched human coders ( r = .54–.65). This article offers a practical guide for evaluating automated coders in psychological testing and discusses practical decisions and ethical considerations.
科研通智能强力驱动
Strongly Powered by AbleSci AI