Customised OCR correction for historical medical text

Customised OCR correction for historical medical text
复制标题

DOI:
10.1109/digitalheritage.2015.7413829
复制
发表时间:
2015-09
期刊:
2015 Digital Heritage
影响因子:
--
通讯作者:
Paul Thompson;J. McNaught;S. Ananiadou
Paul Thompson;J. McNaught;S. Ananiadou
中科院分区:
其他
文献类型:
--
作者:
Paul Thompson;J. McNaught;S. Ananiadou

文献摘要

相似文献

历史文本档案构成了丰富多样的信息来源,由于大规模的数字化努力,这些信息变得越来越容易获取。可搜索访问通常是通过将光学字符识别 (OCR) 软件应用于扫描的页面图像来提供的。然而,自动识别的文本通常包含大量错误,因为 OCR 系统通常经过优化以处理现代文档,并且可能会与历史文档特征(包括可变的打印特征和过时的词汇使用)作斗争。低质量的 OCR 文本会降低历史档案搜索系统的效率,特别是基于复杂文本挖掘 (TM) 技术应用的语义系统。我们报告了针对历史医疗文档定制的新 OCR 校正策略。该方法将基于规则的常规错误纠正与经过医学调整的拼写检查策略相结合,其纠正以待纠正文章出版期间的特定主题语言使用信息为指导。我们的方法的性能优于其他 OCR 后校正策略,将质量较差的文档的字级准确率提高了 16%。
Historical text archives constitute a rich and diverse source of information, which is becoming increasingly readily accessible, owing to large-scale digitisation efforts. Searchable access is typically provided by applying Optical Character Recognition (OCR) software to scanned page images. Often, however, the automatically recognised text contains a large number of errors, since OCR systems are typically optimised to deal with modern documents, and can struggle with historical document features, including variable print characteristics and archaic vocabulary usage. Low quality OCR text can reduce the efficiency of search systems over historical archives, particularly semantic systems that are based on the application of sophisticated text mining (TM) techniques. We report on a new OCR correction strategy, customised for historical medical documents. The method combines rule-based correction of regular errors with a medically-tuned spell-checking strategy, whose corrections are guided by information about subject-specific language usage from the publication period of the article to be corrected. The performance of our method compares favourably to other OCR post-correction strategies, in improving word-level accuracy of poor-quality documents by up to 16%.