Error annotation systems
Error annotation systems
复制标题
错误注释系统
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Hagen Hirschmann
中科院分区:
文献类型:
--
作者:
Anke Lüdeling;Hagen Hirschmann
and says only that this part of the learner utterance is unidiomatic, confl ating an implicit target hypothesis with an error tag (the annotator is only able to know that this expression is unidiomatic if he or she knows a more idiomatic expression). Different target hypotheses are not equivalent; a target hypothesis directly infl uences the following analysis. The Falko corpus consistently has two target hypotheses – the fi rst one deals with clear grammatical errors and the second one also corrects stylistic problems. The need for such an approach becomes clear in (11). The learner utterance in (11) contains a spelling error . The two occurrences of dependance have to be replaced by dependence . From a more abstract perspective, the whole phrase Dependence on gambling sounds unidiomatic if we take into account that the learner wants to refer to a specifi c kind of addiction. Similarly, dependence on drugs appears to be a marked expression as opposed to drug addiction . An annotation that wants to take this into consideration has to separate the description into the annotation of the spelling error and the annotation of the stylistic error in order not to lose one of the pieces of information. Example (12) illustrates this. The examples in this section show how important the step of formulating a target hypothesis is – the subsequent error classifi cation critically depends on this fi rst step. In order to operationalise the fi rst step of the error annotation , one can give guidelines for the formulation of target hypotheses, in addition to the guidelines for assigning error tags, which also need to be evaluated with regard to consistency (see Section 2.6 ). The problem of unclear error identifi cation has been discussed since the beginning of EA. Milton and Chowdhury ( 1994 ) have already suggested that sometimes multiple analyses should be coded in a learner corpus. If (11) Dependance on gambling is something like dependance on drugs (...) (ICLE-CZ-PRAG-0013.3) (12) LU Dependance on gambling TH 1 Dependence on gambling TH 2 Gambling addiction (10) LU it sleeps inside everyone from the start of being TH 1 it sleeps inside everyone since birth TH 2 it sleeps inside everyone from the beginning TH 3 it sleeps inside everyone UNIDIOMATIC 9781107041196c07_p135-158.indd 145 6/11/2015 1:48:09 PM LÜDELING AND HIRSCHMANN 146 the target hypothesis is left implicit or there is only one error analysis , the user is given an error annotation without knowing against which form the utterance was evaluated. In early corpora (pre-multi-layer, pre-XML) it was technically impossible to show the error exponent because errors could only be marked on one token. In corpora that use an XML format it is possible to mark spans, and target hypotheses are sometimes given in the XML mark-up. Only in standoff architectures, however, is it possible to give several competing target hypotheses. Examples of learner corpora with consistent and well-documented (multiple) target hypotheses are the Falko corpus, the trilingual MERLIN corpus (Wisniewski et al. 2013 ) or the Czech as a Second Language corpus (Rosen et al. 2014 ).