Post-Editing Through Approximation and Global Correction

Post-Editing Through Approximation and Global Correction
复制标题

通过近似和全局校正进行后期编辑

DOI:
10.1142/s0218001495000377
复制
发表时间:
1995
影响因子:
1.5
通讯作者:
A. Condit
A. Condit
中科院分区:
计算机科学4区
文献类型:
--
作者:
K. Taghva;J. Borsack;Bryan Bullard;A. Condit

文献摘要

被引文献

相似文献

本文介绍了一种新的自动拼写纠正程序,以处理OCR产生的错误。这里使用的方法基于三个原则:1。在拼写错误和数据库中出现的术语之间进行近似的字符串匹配,而不是整个字典2。从个别文件中获得的本地信息3.使用混淆矩阵,其中包含特定OCR设备所导致的错误本质的固有特定信息。然后,该系统用于处理大约10,000页OCR生成的文档。在通过该算法发现的拼写错误中,约87%被纠正。
This paper describes a new automatic spelling correction program to deal with OCR generated errors. The method used here is based on three principles: 1. Approximate string matching between the misspellings and the terms occuring in the database as opposed to the entire dictionary 2. Local information obtained from the individual documents 3. The use of a confusion matrix, which contains information inherently specific to the nature of errors caused by the particular OCR device This system is then utilized to process approximately 10,000 pages of OCR generated documents. Among the misspellings discovered by this algorithm, about 87% were corrected.