摘要
This study investigated whether ranking-based translation quality assessment provides more reliable, interpretable, and decision-relevant evidence than traditional rating-based evaluation. Using an English-to-Persian empirical approach, 72 bilingual raters assessed 120 source-text segments across six domains and five candidate conditions: professional human translation, post-edited machine translation, neural machine translation, LLM-generated translation, and controlled-error translation. The research compared various methods, including analytic ratings, pairwise comparisons, tie-enabled judgments, best – worst scaling, side-by-side judgments, and short rank ordering. Psychometric analyses employed many-facet Rasch modeling, Bradley – Terry, Davidson, Plackett – Luce, Bayesian hierarchical comparison models, generalizability analysis, and fairness diagnostics. Results indicated that ranking-based models yielded stronger reliability, generalizability, classification consistency, and rank stability than unadjusted ratings. Among these, professional human translation had the highest latent quality estimate, followed by post-edited machine translation, LLM-generated translation, neural machine translation, and controlled-error translation. However, the difference between human and post-edited machine translation was not decisive under the preset probability threshold, suggesting practical proximity. Ratings remained useful for diagnosing specific weaknesses, especially regarding adequacy and terminology. The findings support a hybrid, uncertainty-aware framework in which ratings explain translation quality, while rankings enhance comparative decision-making in modern MT and LLM evaluation across domains. This approach boosts validity, transparency, and fairness in high-stakes translation assessments.