摘要
This study aims to revisit the Menzerath-Altmann Law (MAL) at the sentence-clause-word level and the paragraph-sentence-clause level across nine languages from the Indo-European language family: Albanian (Albanian), Armenian (Armenic), Lithuanian (Balto-Slavic), Gaelic (Celtic), Danish (North Germanic), German (West Germanic), Urdu (Indo-Iranian), Italian (Italian Romance), and French (Western Romance), and across four registers: press, general prose, academic prose, and fiction. The study is based on nine one-million-token Stanza-annotated multilingual balanced corpora that followed the Brown Corpus sampling frame, which made the corpora highly comparable. The results show that the frequency distributions of the sentence and paragraph lengths are strongly right-skewed, regardless of language and register. At the sentence-clause-word level, MAL holds robustly across the nine Indo-European languages (R2 = 0.907 to 0.990). Nevertheless, at the paragraph–sentence–clause level, MAL is not consistently valid: only Albanian and Gaelic have R2 ≥ 0.7. MAL fits are register-sensitive. At the sentence–clause–word level, the MAL fit of fiction is significantly weaker than those of press, general prose, and academic prose. At the paragraph–sentence–clause level, only press shows moderately good fits (e.g., Danish, Gaelic, Lithuanian). The study also shows that MAL’s cross-linguistic differences increase and goodness-of-fit decreases from sentence–clause–word level to paragraph–sentence–clause level.