The Past is Not a Foreign Country: Detecting Semantically Similar Terms across Time

The Past is Not a Foreign Country: Detecting Semantically Similar Terms across Time
复制标题

过去不是外国:跨时间检测语义相似的术语

DOI:
10.1109/tkde.2016.2591008
复制
发表时间:
2016
期刊:
IEEE Trans. Knowl. Data Eng.
影响因子:
--
通讯作者:
Katsumi Tanaka
Katsumi Tanaka
中科院分区:
--
文献类型:
--
作者:
Yating Zhang;Adam Jatowt;Sourav S. Bhowmick;Katsumi Tanaka

文献摘要

相似文献

由于大规模的数字化和保存工作,许多过去文件的档案和收藏最近已经可以使用。图书馆、国家档案馆和其他记忆机构已经开始向感兴趣的用户开放他们的藏品。然而,由于不同的背景和过去的语言,在这样的集合中进行搜索通常需要了解适当的关键字。因此,非专业用户可能难以概念化合适的查询,因为他们对过去的知识通常是有限的。在本文中,我们提出了一种新的时间对应检测任务,该任务需要在过去找到在语义上与给定输入当前词最接近的词。我们提出的方法是基于向量空间变换,将现在的分布式单词表示映射到过去的分布式单词表示。这种方法的关键问题是获得正确的训练集,这些训练集可以用于各种不同的文档集合和任意时间段。为了解决这个问题,我们提出了一种有效的技术,用于自动构建用于寻找转换的术语的种子对。我们测试了所提出的方法在短时间段和长时间段(例如100年)的性能。我们的实验表明,当查询的字面形式与其时间对应的查询具有不同的字面形式时,所提出的方法在纽约时报注释语料库上的性能平均高出113%,在MRR中的时报档案上平均高出28%。
Numerous archives and collections of past documents have become available recently thanks to mass scale digitization and preservation efforts. Libraries, national archives, and other memory institutions have started opening up their collections to interested users. Yet, searching within such collections usually requires knowledge of appropriate keywords due to different context and language of the past. Thus, non-professional users may have difficulties with conceptualizing suitable queries, as, typically, their knowledge of the past is limited. In this paper, we propose a novel approach for the temporal correspondence detection task that requires finding terms in the past which are semantically closest to a given input present term. The approach we propose is based on vector space transformation that maps the distributed word representation in the present to the one in the past. The key problem in this approach is obtaining correct training set that could be used for a variety of diverse document collections and arbitrary time periods. To solve this problem, we propose an effective technique for automatically constructing seed pairs of terms to be used for finding the transformation. We test the performance of proposed approaches over short as well as long time frames such as 100 years. Our experiments demonstrate that the proposed methods outperform the best-performing baseline by 113 percent for the New York Times Annotated Corpus and by 28 percent for the Times Archive in MRR on average, when the query has a different literal form from its temporal counterpart.