可比性
考试(生物学)
拉什模型
任务(项目管理)
心理学
可靠性(半导体)
集合(抽象数据类型)
认知心理学
变化(天文学)
出声思维法
自然语言处理
数学教育
计算机科学
发展心理学
数学
管理
程序设计语言
可用性
人机交互
量子力学
生物
功率(物理)
组合数学
物理
古生物学
经济
天体物理学
作者
Cyril J. Weir,Jessica R. W. Wu
出处
期刊:Language Testing
[SAGE Publishing]
日期:2006-03-25
卷期号:23 (2): 167-197
被引量:47
标识
DOI:10.1191/0265532206lt326oa
摘要
Examination boards are often criticized for their failure to provide evidence of comparability across forms, and few such studies are publicly available. This study aims to investigate the extent to which three forms of the General English Proficiency Test Intermediate Speaking Test (GEPTS-I) are parallel in terms of two types of validity evidence: parallel-forms reliability and content validity. The three trial test forms, each containing three different task types (read-aloud, answering questions and picture description), were administered to 120 intermediate-level EFL learners in Taiwan. The performance data from the different test forms were analysed using classical procedures and Multi-Faceted Rasch Measurement (MFRM). Various checklists were also employed to compare the tasks in different forms qualitatively in terms of content. The results showed that all three test forms were statistically parallel overall and Forms 2 and 3 could also be considered parallel at the individual task level. Moreover, sources of variation to account for the variable difficulty of tasks in Form 1 were identified by the checklists. Results of the study provide insights for further improvement in parallel-form reliability of the GEPTS-I at the task level and offer a set of methodological procedures for other exam boards to consider.
科研通智能强力驱动
Strongly Powered by AbleSci AI