专题地图
名词
自然语言处理
人工智能
计算机科学
聚类分析
语言学
主题结构
抓住
层次聚类
特征(语言学)
主题分析
点(几何)
数学
地理
定性研究
地图学
社会学
社会科学
哲学
程序设计语言
几何学
标识
DOI:10.1080/09296174.2017.1339441
摘要
Generally, human brains can grasp intuitively the gist of thematic content of different texts through comprehensive reading, and such human-like generalization process may be accomplished with a more exact basis. With three representative text types in Chinese and English from two comparative corpora as our focus, that is, LCMC (the Lancaster Corpus of Mandarin Chinese) and Frown (the Freiburg-Brown Corpus of American English), this study compares thematic characteristics of these texts with PAM (Partition around Medoids) and HA (Hierarchical Agglomerative) clustering via three quantitative indicators, namely, TC (Thematic Concentration), STC (Secondary Thematic Concentration) and PTC (Proportional Thematic Concentration). The results show that: (1) eigenvectors standing for the thematic characteristic of three text types can be clustered into their corresponding categories in both Chinese and English; (2) two contributing factors are identified for the clustering results. One is the differences of TC, STC and PTC values of three text types lying in different hierarchical levels; the other is the differences of the percentages of 'thematic words', especially nouns at the pre-h-point and pre-2 h-point domain in three text types. The characterization of three text types as thematic-intensive (Official Document), thematic-balanced (News) and thematic-dispersive (Fiction) bears a cross-linguistic similarity in both Chinese and English.
科研通智能强力驱动
Strongly Powered by AbleSci AI