Computer detection of typographical errors

Computer detection of typographical errors
复制标题

计算机检测印刷错误

DOI:
10.1109/tpc.1975.6593963
复制
发表时间:
1975
影响因子:
1.7
通讯作者:
Cherry
Cherry
中科院分区:
人文科学4区
文献类型:
--
作者:
R. Morris;Lorinda;Cherry

文献摘要

被引文献

相似文献

描述为 UNIX 分时系统编写的计算机程序,该程序将在文档中查找包含印刷错误的单词的任务减少几个数量级。该程序是自适应的,因为它使用文档本身的统计数据进行分析。在第一次浏览该文档时,准备了图表和卦词频率表。第二次浏览文档时,会分解各个单词,并将每个单词中的图表和三元组与表中的频率进行比较。每个单词都有一个索引,该索引反映了这样的假设:给定单词中的三元组是从生成三元组表的同一来源生成的。这些单词按照其索引的降序排列并打印。禁止打印 2726 个常见技术英语单词表中出现的单词。
Describes a computer program written for the UNIX time-sharing system which reduces by several orders of magnitude the task of finding words in a document which contain typographical errors. The program is adaptive in the sense that it uses statistics from the document itself for its analysis. In a first pass through the document, a table of diagram and trigram frequencies is prepared. The second pass through the document breaks out individual words and compares the diagrams and trigrams in each word with the frequencies from the table. An index is given to each word which reflects the hypothesis that the trigrams in the given word were produced from the same source that produced the trigram table. The words are sorted in decreasing order of their indices and printed. Printing is suppressed for words appearing in a table of 2726 common technical English words.