雅卡索引
概率逻辑
规范化(社会学)
相似性(几何)
度量(数据仓库)
数据挖掘
计算机科学
相似性度量
集合(抽象数据类型)
余弦相似度
索引(排版)
数学
人工智能
模式识别(心理学)
万维网
人类学
图像(数学)
程序设计语言
社会学
作者
Nees Jan van Eck,Ludo Waltman
摘要
Abstract In scientometric research, the use of cooccurrence data is very common. In many cases, a similarity measure is employed to normalize the data. However, there is no consensus among researchers on which similarity measure is most appropriate for normalization purposes. In this article, we theoretically analyze the properties of similarity measures for cooccurrence data, focusing in particular on four well‐known measures: the association strength, the cosine, the inclusion index, and the Jaccard index. We also study the behavior of these measures empirically. Our analysis reveals that there exist two fundamentally different types of similarity measures, namely, set‐theoretic measures and probabilistic measures. The association strength is a probabilistic measure, while the cosine, the inclusion index, and the Jaccard index are set‐theoretic measures. Both our theoretical and our empirical results indicate that cooccurrence data can best be normalized using a probabilistic measure. This provides strong support for the use of the association strength in scientometric research.
科研通智能强力驱动
Strongly Powered by AbleSci AI