Measuring the Correctness of Double-Keying: Error Classification and Quality Control in a Large Corpus of TEI-Annotated Historical Text

Measuring the Correctness of Double-Keying: Error Classification and Quality Control in a Large Corpus of TEI-Annotated Historical Text
复制标题

测量双重键入的正确性:TEI 注释历史文本大型语料库中的错误分类和质量控制

DOI:
10.4000/jtei.739
复制
发表时间:
2013
期刊:
Journal of Surgery and Research
影响因子:
--
通讯作者:
Alexander Geyken
Alexander Geyken
中科院分区:
--
文献类型:
--
作者:
S. Haaf;Frank Wiegand;Alexander Geyken

文献摘要

被引文献

相似文献

在大规模数字化方法中,双键被认为是错误率最低的方法。这种方法需要由两个不同的操作员对文本进行两次独立的转录。它特别适合于历史文本,因为历史文本经常表现出不足之处,如糟糕的原版文本,或其他困难,如拼写变化或复杂的文本结构。使用双键方法的数据输入服务提供商通常宣传非常高的准确率(约99.95%至99.98%)。这些广告百分比通常是基于小样本来估计的,并且很少涉及实际文本数量或已经校对的文本体裁、关于错误类型、校对者等。为了获得关于该问题的重要数据,有必要分析代表不同文本类型的平衡样本的大量文本,以区分结构XML/TEI水平和排版水平,并且区分可能源自不同来源且可能不是同样严重的各种类型的错误。本文介绍了一种广泛而复杂的分析和纠正双键错误的方法,该方法已被DFG资助的项目德国文本档案馆(以下简称DTA)所应用,以评估并更好地提高双键DTA文本的转录和注释的准确性。给出了对大量文本校对结果的统计分析,验证了双键方法的常见准确率。
Among mass digitization methods, double-keying is considered to be the one with the lowest error rate. This method requires two independent transcriptions of a text by two different operators. It is particularly well suited to historical texts, which often exhibit deficiencies like poor master copies or other difficulties such as spelling variation or complex text structures. Providers of data entry services using the double-keying method generally advertise very high accuracy rates (around 99.95% to 99.98%). These advertised percentages are generally estimated on the basis of small samples, and little if anything is said about either the actual amount of text or the text genres which have been proofread, about error types, proofreaders, etc. In order to obtain significant data on this problem it is necessary to analyze a large amount of text representing a balanced sample of different text types, to distinguish the structural XML/TEI level from the typographical level, and to differentiate between various types of errors which may originate from different sources and may not be equally severe. This paper presents an extensive and complex approach to the analysis and correction of double-keying errors which has been applied by the DFG-funded project "Deutsches Textarchiv" (German Text Archive, hereafter DTA) in order to evaluate and preferably to increase the transcription and annotation accuracy of double-keyed DTA texts. Statistical analyses of the results gained from proofreading a large quantity of text are presented, which verify the common accuracy rates for the double-keying method.