潜在Dirichlet分配
计算机科学
主题模型
公制(单位)
潜在语义分析
情报检索
线性子空间
潜变量
降维
可解释性
人工智能
数据科学
自然语言处理
数学
经济
运营管理
几何学
作者
David Mimno,Hanna Wallach,Edmund M. Talley,Miriam Leenders,Andrew McCallum
出处
期刊:University of Massachusetts Amherst - ScholarWorks@UMassAmherst
日期:2011-07-27
卷期号:: 262-272
被引量:1248
摘要
Large organizations often face the critical challenge of sharing information and maintaining connections between disparate subunits. Tools for automated analysis of document collections, such as topic models, can provide an important means for communication. The value of topic modeling is in its ability to discover interpretable, coherent themes from unstructured document sets, yet it is not unusual to find semantic mismatches that substantially reduce user confidence. In this paper, we first present an expert-driven topic annotation study, undertaken in order to obtain an annotated set of baseline topics and their distinguishing characteristics. We then present a metric for detecting poor-quality topics that does not rely on human feedback or external reference corpora. Next we introduce a new topic model that incorporates salient properties of this metric. We show significant gains in topic quality on a substantial document collection from the National Institutes of Health, measured using both automated evaluation metrics and expert evaluations.
科研通智能强力驱动
Strongly Powered by AbleSci AI