Towards information retrieval on historical document collections: the role of matching procedures and special lexica

Towards information retrieval on historical document collections: the role of matching procedures and special lexica
复制标题

历史文献收藏的信息检索:匹配程序和特殊词汇的作用

DOI:
--
复制
发表时间:
2011
影响因子:
2.3
通讯作者:
Andreas W. Neumann
Andreas W. Neumann
中科院分区:
计算机科学4区
文献类型:
--
作者:
Annette Gotscharek;Ulrich Reffle;Christoph Ringlstetter;K. Schulz;Andreas W. Neumann

文献摘要

被引文献

相似文献

由于历史文献中存在大量的拼写变体,标准的信息检索方法在历史文献检索中往往不能取得令人满意的结果。为了提高搜索引擎的查全率,查询中使用的现代单词必须与文档中找到的相应历史变体相关联。在文献中,使用(1)特殊的匹配程序和(2)词汇的历史语言已被建议作为两种替代方法来解决这个问题。在本文的第一部分中,我们展示了匹配程序和词汇的构建如何相互受益,从而将这两种方法结合起来。提出了一种基于语料库分析的匹配规则和历史词典交叉构建的工具。在本文的第二部分,我们问,如果匹配程序本身就足以解除IR的历史文本到一个令人满意的水平。由于历史语言在几个世纪中不断变化,要得到答案并不容易。我们目前的实验中,从四个世纪的文本集合中的匹配过程的性能进行了研究。在对遗漏词汇进行分类后,我们测量了每个阶段匹配过程的精确度和召回率。结果表明,对于早期,单独的匹配程序不会导致令人满意的结果。然后,我们描述的实验中获得的不同大小的历史词汇的召回估计。
Due to the large number of spelling variants found in historical texts, standard methods of Information Retrieval (IR) fail to produce satisfactory results on historical document collections. In order to improve recall for search engines, modern words used in queries have to be associated with corresponding historical variants found in the documents. In the literature, the use of (1) special matching procedures and (2) lexica for historical language have been suggested as two alternative ways to solve this problem. In the first part of the paper, we show how the construction of matching procedures and lexica may benefit from each other, leading the way to a combination of both approaches. A tool is presented where matching rules and a historical lexicon are built in an interleaved way based on corpus analysis. In the second part of the paper, we ask if matching procedures alone suffice to lift IR on historical texts to a satisfactory level. Since historical language changes over centuries, it is not simple to obtain an answer. We present experiments where the performance of matching procedures in text collections from four centuries is studied. After classifying missed vocabulary, we measure precision and recall of the matching procedure for each period. Results indicate that for earlier periods, matching procedures alone do not lead to satisfactory results. We then describe experiments where the gain for recall obtained from historical lexica of distinct sizes is estimated.