News Topic Tracking and Re-ranking with Query Expansion Based on Near-Duplicate Detection

News Topic Tracking and Re-ranking with Query Expansion Based on Near-Duplicate Detection
复制标题

DOI:
10.1007/978-3-642-10467-1_66
复制
发表时间:
2009-12
期刊:
--
影响因子:
--
通讯作者:
Xiaomeng Wu;I. Ide;S. Satoh
Xiaomeng Wu;I. Ide;S. Satoh
中科院分区:
其他
文献类型:
--
作者:
Xiaomeng Wu;I. Ide;S. Satoh

文献摘要

被引文献

相似文献

数字存储容量的增加使得大规模新闻视频档案的创建成为可能。要充分利用新闻档案,就必须把握新闻故事的发展脉络和依赖关系。考虑到这个问题,我们研究了新闻故事的跟踪和重新排名方法。作为试验台的档案包括3万多篇新闻报道。本文提出了一种基于基于文本的近重复的查询扩展算法挖掘主题相关故事的新方案。实验表明,基于近重复约束的查询扩展算法优于仅使用文本特征的传统方法。
Increase of digital storage capacity enabled the creation of large-scale news video archives. To make full use of the archive, it is necessary to grasp the development and dependencies of news stories. Considering this problem, we investigate tracking and re-ranking methodologies of news stories. The archive used as a test-bed consists of more than 30,000 news stories. This paper proposes a novel scheme of mining topic-related stories through a query-expansion algorithm on the basis of near duplicates built on top of text. Experiments showed that the query-expansion algorithm based on near-duplicate constraints outperformed traditional methods that only use textual features.