Error annotation systems

Error annotation systems
复制标题

错误注释系统

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Hagen Hirschmann
Hagen Hirschmann
中科院分区:
--
文献类型:
--
作者:
Anke Lüdeling;Hagen Hirschmann

文献摘要

被引文献

相似文献

并仅表示学习者话语的这一部分不惯用,将隐式目标假设与错误标签混为一谈(注释者只有在知道更惯用的表达时才能够知道该表达不惯用)。不同的目标假设并不等价;目标假设直接影响接下来的分析。 Falko 语料库始终有两个目标假设——第一个假设处理明显的语法错误,第二个假设也纠正文体问题。这种方法的必要性在(11)中变得很清楚。 (11) 中的学习者话语包含拼写错误 。两次出现的依赖必须用依赖来代替。从更抽象的角度来看,如果我们考虑到学习者想要指的是一种特定类型的成瘾,那么整个短语“赌博依赖”听起来不惯用。同样,与毒瘾相反,对药物的依赖似乎是一个明显的表现。想要考虑到这一点的注释必须将描述分为拼写错误的注释和文体错误的注释,以免丢失其中一条信息。例(12)说明了这一点。本节中的示例显示了制定目标假设的步骤是多么重要——随后的错误分类关键取决于这第一步。为了实施错误注释的第一步,除了分配错误标签的指南之外,还可以给出制定目标假设的指南,还需要评估一致性(参见第2.6节)。错误识别不明确的问题从 EA 诞生之初就一直在讨论。 Milton 和 Chowdhury(1994)已经建议有时应该在学习者语料库中编码多种分析。如果 (11) 对赌博的依赖就像对毒品的依赖 (...) (ICLE-CZ-PRAG-0013.3) (12) LU 对赌博的依赖 TH 1 对赌博的依赖 TH 2 赌博成瘾 (10) LU 从一开始就沉睡在每个人的内心 TH 1 从出生起就沉睡在每个人的内心 TH 2 从一开始就沉睡在每个人的内心 TH 3 它沉睡在每个人的内心 UNIDIOMATIC 9781107041196c07_p135-158.indd 145 6/11/2015 1:48:09 PM LÜDELING AND HIRSCHMANN 146 目标假设是隐含的或者只有一个错误分析,用户会得到一个错误注释,而不知道根据哪种形式评估话语。在早期语料库(多层之前、XML 之前)中,从技术上来说,显示错误指数是不可能的,因为错误只能标记在一个标记上。在使用 XML 格式的语料库中,可以标记跨度,并且目标假设有时会在 XML 标记中给出。然而,只有在对峙架构中,才有可能给出几个相互竞争的目标假设。具有一致且记录充分的(多个)目标假设的学习者语料库的示例包括 Falko 语料库、三语 MERLIN 语料库(Wisniewski 等人,2013 年)或捷克语作为第二语言语料库(Rosen 等人,2014 年)。
and says only that this part of the learner utterance is unidiomatic, confl ating an implicit target hypothesis with an error tag (the annotator is only able to know that this expression is unidiomatic if he or she knows a more idiomatic expression). Different target hypotheses are not equivalent; a target hypothesis directly infl uences the following analysis. The Falko corpus consistently has two target hypotheses – the fi rst one deals with clear grammatical errors and the second one also corrects stylistic problems. The need for such an approach becomes clear in (11). The learner utterance in (11) contains a spelling error . The two occurrences of dependance have to be replaced by dependence . From a more abstract perspective, the whole phrase Dependence on gambling sounds unidiomatic if we take into account that the learner wants to refer to a specifi c kind of addiction. Similarly, dependence on drugs appears to be a marked expression as opposed to drug addiction . An annotation that wants to take this into consideration has to separate the description into the annotation of the spelling error and the annotation of the stylistic error in order not to lose one of the pieces of information. Example (12) illustrates this. The examples in this section show how important the step of formulating a target hypothesis is – the subsequent error classifi cation critically depends on this fi rst step. In order to operationalise the fi rst step of the error annotation , one can give guidelines for the formulation of target hypotheses, in addition to the guidelines for assigning error tags, which also need to be evaluated with regard to consistency (see Section 2.6 ). The problem of unclear error identifi cation has been discussed since the beginning of EA. Milton and Chowdhury ( 1994 ) have already suggested that sometimes multiple analyses should be coded in a learner corpus. If (11) Dependance on gambling is something like dependance on drugs (...) (ICLE-CZ-PRAG-0013.3) (12) LU Dependance on gambling TH 1 Dependence on gambling TH 2 Gambling addiction (10) LU it sleeps inside everyone from the start of being TH 1 it sleeps inside everyone since birth TH 2 it sleeps inside everyone from the beginning TH 3 it sleeps inside everyone UNIDIOMATIC 9781107041196c07_p135-158.indd 145 6/11/2015 1:48:09 PM LÜDELING AND HIRSCHMANN 146 the target hypothesis is left implicit or there is only one error analysis , the user is given an error annotation without knowing against which form the utterance was evaluated. In early corpora (pre-multi-layer, pre-XML) it was technically impossible to show the error exponent because errors could only be marked on one token. In corpora that use an XML format it is possible to mark spans, and target hypotheses are sometimes given in the XML mark-up. Only in standoff architectures, however, is it possible to give several competing target hypotheses. Examples of learner corpora with consistent and well-documented (multiple) target hypotheses are the Falko corpus, the trilingual MERLIN corpus (Wisniewski et al. 2013 ) or the Czech as a Second Language corpus (Rosen et al. 2014 ).