Because of the increasing popularity and use of performance testing and rating scales in ESL writing, many scholars have stressed the need to pay close attention to the raters themselves. Hence, a number of studies have been explored to determine the impact of rater characteristics, rating method, and rating patterns on interrater agreement. There is, however, little research that attempted to analyze interrater variability in a more micro-level and to compare raters’ specific comments to determine whether they arrive at the same rating for the same reasons.
Specifically, this paper sought to determine:
(1) the level of agreement among raters and the factors that might have influenced the results,
(2) the similarities of raters’ reasons for arriving at the same rating, and
(3) the rater-related factors that might account for their differences in assessing ESL learners’ written production.
Three groups of experienced raters rated 39 essays and provided reasons for arriving at such ratings. Using a mixed method approach, findings revealed that raters posted a fair interrater agreement and that they had different reasons for arriving at the same rating. Such findings were mainly attributed to the raters’ different scoring focus and rating scale. This study has implications for ESL writing assessment practices and for future studies.