Incorporating terminology evolution for query translation in text retrieval with association rules

Incorporating terminology evolution for query translation in text retrieval with association rules
复制标题

将文本检索中查询翻译的术语演变与关联规则结合起来

DOI:
10.1145/1871437.1871730
复制
发表时间:
2010
期刊:
Proceedings of the 19th ACM international conference on Information and knowledge management
影响因子:
--
通讯作者:
Anna Feldman
Anna Feldman
中科院分区:
--
文献类型:
--
作者:
A. Kaluarachchi;A. Varde;Srikanta J. Bedathur;G. Weikum;Jing Peng;Anna Feldman

文献摘要

被引文献

相似文献

带有时间戳的文档(例如新闻专线文章、博客文章和其他网页)通常在线存档。当这些档案涵盖很长的时间跨度时,其中的术语可能会发生重大变化。因此,当用户在这些文档上提出与历史信息有关的查询时,需要考虑这些时间变化来翻译查询,以向用户提供准确的响应。例如,对斯里兰卡的查询应该自动检索具有其以前名称Ceylon的文档。我们称这种概念为SITAC,即,语义相同的临时改变概念。为了发现SITACs,我们提出了一种基于一个新的框架构成的自然语言处理,关联规则挖掘,上下文相似性作为一种学习技术的集成方法。所提出的方法已与真实的数据进行了实验,并已被发现产生良好的效果方面的效率和准确性。
Time-stamped documents such as newswire articles, blog posts and other web-pages are often archived online. When these archives cover long spans of time, the terminology within them could undergo significant changes. Hence, when users pose queries pertaining to historical information, over such documents, the queries need to be translated, taking into account these temporal changes, to provide accurate responses to users. For example, a query on Sri Lanka should automatically retrieve documents with its former name Ceylon. We call such concepts SITACs, i.e., Semantically Identical Temporally Altering Concepts. In order to discover SITACs, we propose an approach based on a novel framework constituting an integration of natural language processing, association rule mining, and contextual similarity as a learning technique. The proposed approach has been experimented with real data and has been found to yield good results with respect to efficiency and accuracy.