Enabling information retrieval on historical document collections: the role of matching procedures and special lexica

Enabling information retrieval on historical document collections: the role of matching procedures and special lexica
复制标题

启用历史文档集合的信息检索:匹配程序和特殊词汇的作用

DOI:
--
复制
发表时间:
2009
期刊:
Workshop on Analytics for Noisy Unstructured Text Data
影响因子:
--
通讯作者:
K. Schulz
K. Schulz
中科院分区:
--
文献类型:
--
作者:
Annette Gotscharek;Andreas W. Neumann;Ulrich Reffle;Christoph Ringlstetter;K. Schulz

文献摘要

被引文献

相似文献

由于历史文献中存在大量的拼写变体,标准的信息检索方法在历史文献检索中往往不能取得令人满意的结果。为了提高搜索引擎的查全率,查询中使用的现代单词必须与文档中找到的相应历史变体相关联。在文献中,使用(1)特殊的匹配程序和(2)词汇的历史语言已被建议作为两种方法来解决这个问题。在本文的第一部分中,我们展示了如何建设的匹配程序和词汇可能会受益于对方,导致这两种方法的结合。提出了一种基于语料库分析的匹配规则和历史词典交叉构建的工具。本文的第二部分考虑的一个关键问题是,如果匹配程序本身就足以解除IR的历史文本到一个令人满意的水平。由于历史语言在几个世纪中不断变化,要得到答案并不容易。我们目前的实验中,从四个世纪的文本集合中的匹配过程的性能进行了研究。在对遗漏词汇进行分类后,我们测量了每个阶段匹配过程的精确度和召回率。我们的研究结果表明,早期的历史词汇是一个重要的纠正匹配程序在IR应用程序。
Due to the large number of spelling variants found in historical texts, standard methods of Information Retrieval (IR) fail to produce satisfactory results on historical document collections. In order to improve recall for search engines, modern words used in queries have to be associated with corresponding historical variants found in the documents. In the literature, the use of (1) special matching procedures and (2) lexica for historical language have been suggested as two ways to solve this problem. In the first part of the paper we show how the construction of matching procedures and lexica may benefit from each other, leading the way to a combination of both approaches. A tool is presented where matching rules and a historical lexicon are built in an interleaved way based on corpus analysis. A crucial question considered in the second part of the paper is if matching procedures alone suffice to lift IR on historical texts to a satisfactory level. Since historical language changes over centuries it is not simple to obtain an answer. We present experiments where the performance of matching procedures in text collections from four centuries is studied. After classifying missed vocabulary, we measure precision and recall of the matching procedure for each period. Our results indicate that for earlier periods historical lexica represent an important corrective to matching procedures in IR applications.