Digitised historical text: Does it have to be mediOCRe?

Digitised historical text: Does it have to be mediOCRe?
复制标题

数字化历史文本:它一定是平庸的吗?

DOI:
--
复制
发表时间:
2012
期刊:
Conference on Natural Language Processing
影响因子:
--
通讯作者:
R. Tobin
R. Tobin
中科院分区:
--
文献类型:
--
作者:
Beatrice Alex;Claire Grover;Ewan Klein;R. Tobin

文献摘要

被引文献

相似文献

本文报道了作为文本挖掘的第一步,提高历史文本的光学字符识别(ocr)质量的实验。我们分析了ocred文本的质量相比,一个金标准,并展示了如何可以通过执行两个自动校正步骤来提高。我们还展示了影响,这可能会对命名实体识别在一个初步的外部评估。这项工作是作为贸易后果项目的一部分进行的,该项目的重点是对历史文件进行文本挖掘,以研究世纪大英帝国的贸易。
This paper reports on experiments to improve the Optical Character Recognition (ocr) quality of historical text as a preliminary step in text mining. We analyse the quality of ocred text compared to a gold standard and show how it can be improved by performing two automatic correction steps. We also demonstrate the impact this can have on named entity recognition in a preliminary extrinsic evaluation. This work was performed as part of the Trading Consequences project which is focussed on text mining of historical documents for the study of nineteenth century trade in the British Empire.