Enhancing cluster labeling using wikipedia

Enhancing cluster labeling using wikipedia
复制标题

DOI:
10.1145/1571941.1571967
复制
发表时间:
2009-07
期刊:
Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval
影响因子:
--
通讯作者:
David Carmel;Haggai Roitman;Naama Zwerdling
David Carmel;Haggai Roitman;Naama Zwerdling
中科院分区:
其他
文献类型:
--
作者:
David Carmel;Haggai Roitman;Naama Zwerdling

文献摘要

被引文献

相似文献

这项工作探讨利用维基百科,免费的在线百科全书的聚类标记增强。我们描述了一个一般框架,聚类标签,提取候选标签从维基百科除了重要的条款,直接从文本中提取。每个候选人的“标签质量”然后由几个独立的法官和评价最高的候选人被推荐标签。我们的实验结果表明,维基百科的标签同意手动标签由人类相关联的集群,远远超过直接从文本中提取的重要条款。我们发现,在大多数情况下,即使当人类的相关标签出现在文本中,纯统计方法很难识别它们作为良好的描述符。此外,我们的实验表明,在我们的测试集合中,超过85%的聚类,手动标签(或变形,或它的同义词)出现在我们的系统推荐的前五个标签。
This work investigates cluster labeling enhancement by utilizing Wikipedia, the free on-line encyclopedia. We describe a general framework for cluster labeling that extracts candidate labels from Wikipedia in addition to important terms that are extracted directly from the text. The "labeling quality" of each candidate is then evaluated by several independent judges and the top evaluated candidates are recommended for labeling. Our experimental results reveal that the Wikipedia labels agree with manual labels associated by humans to a cluster, much more than with significant terms that are extracted directly from the text. We show that in most cases even when human's associated label appears in the text, pure statistical methods have difficulty in identifying them as good descriptors. Furthermore, our experiments show that for more than 85% of the clusters in our test collection, the manual label (or an inflection, or a synonym of it) appears in the top five labels recommended by our system.