Evaluating the Impact of OCR Errors on Topic Modeling

Evaluating the Impact of OCR Errors on Topic Modeling
复制标题

评估 OCR 错误对主题建模的影响

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Asian Digital Libraries
影响因子:
--
通讯作者:
A. Jatowt
A. Jatowt
中科院分区:
--
文献类型:
--
作者:
Stephen Mutuvi;A. Doucet;Moses Odeo;A. Jatowt

文献摘要

被引文献

相似文献

由于各种原因,历史文档给字符识别带来了挑战,例如不同材料之间的字体差异、缺乏拼写相同单词的拼写不同的拼写标准、材料质量以及无法获得已知历史拼写变体的词典。因此,这些文档的光学字符识别(OCR)通常产生令人不满意的OCR准确性,使数字材料只能部分地被发现,并且它们持有的数据难以处理。在本文中,我们探索了OCR错误对从包含历史OCR文档文本的语料库中识别主题的影响。基于在OCR文本语料库上进行的实验,我们观察到OCR噪声对主题建模算法生成的主题的稳定性和连贯性产生负面影响,并量化了这种影响的强度。
Historical documents pose a challenge for character recognition due to various reasons such as font disparities across different materials, lack of orthographic standards where same words are spelled differently, material quality and unavailability of lexicons of known historical spelling variants. As a result, optical character recognition (OCR) of those documents often yield unsatisfactory OCR accuracy and render digital material only partially discoverable and the data they hold difficult to process. In this paper, we explore the impact of OCR errors on the identification of topics from a corpus comprising text from historical OCRed documents. Based on experiments performed on OCR text corpora, we observe that OCR noise negatively impacts the stability and coherence of topics generated by topic modeling algorithms and we quantify the strength of this impact.