计算机科学
自然语言处理
人工智能
句法结构
语法
标识
DOI:10.1080/09296174.2025.2569581
摘要
This study addresses the calculation and evaluation of syntactic distance, which is a quantitative measure of structural similarity or divergence between languages. Building on existing alignment-based, feature-based and data-driven approaches, we introduce a novel hypergraph-based metric that assesses syntactic distance through structural alignment while explicitly incorporating word order features. The approach is then applied to a multilingual parallel corpus annotated within the Universal Dependencies (UD) framework, yielding syntactic distances between English and 19 non-English languages. Empirical evaluation further demonstrates the robustness and effectiveness of the proposed measure. Compared with approaches that ablate the hypergraph formalism, ignore word order or rely solely on data-driven metrics, the new metric proves robust under random sampling variation and effectively captures syntactic distance: statistical analyses show that intra-group language pairs exhibit significantly shorter syntactic distances than inter-group pairs. This approach thus provides a novel, formally grounded perspective on language distance based purely on structural properties.
科研通智能强力驱动
Strongly Powered by AbleSci AI