计算机科学
主题模型
相似性(几何)
分歧(语言学)
聚类分析
推论
光学(聚焦)
引用
数据科学
情报检索
词(群论)
人工智能
万维网
数学
语言学
物理
哲学
光学
图像(数学)
几何学
作者
Shuo Xu,Dongsheng Zhai,Feifei Wang,Xin An,Hongshen Pang,Yirong Sun
摘要
It is increasingly important to build topic linkages between scientific publications and patents for the purpose of understanding the relationships between science and technology. Previous studies on the linkages mainly focus on the analysis of nonpatent references on the front page of patents, or the resulting citation‐link networks, but with unsatisfactory performance. In the meanwhile, abundant mentioned entities in the scholarly articles and patents further complicate topic linkages. To deal with this situation, a novel statistical entity‐topic model (named the CCorrLDA2 model), armed with the collapsed Gibbs sampling inference algorithm, is proposed to discover the hidden topics respectively from the academic articles and patents. In order to reduce the negative impact on topic similarity calculation, word tokens and entity mentions are grouped by the Brown clustering method. Then a topic linkages construction problem is transformed into the well‐known optimal transportation problem after topic similarity is calculated on the basis of symmetrized Kullback–Leibler (KL) divergence. Extensive experimental results indicate that our approach is feasible to build topic linkages with more superior performance than the counterparts.
科研通智能强力驱动
Strongly Powered by AbleSci AI