Supporting the Cognitive Process in Annotation Tasks

Supporting the Cognitive Process in Annotation Tasks
复制标题

支持注释任务中的认知过程

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Eyke Hüllermeier
Eyke Hüllermeier
中科院分区:
--
文献类型:
--
作者:
Nina Seemann;Michaela Geierhos;Marie;Doris Tophinke;Marcel Wever;Eyke Hüllermeier

文献摘要

被引文献

相似文献

注释工具通常使用公共文本分析流水线,其中(I)进行标记化,(Ii)检测句尾,(Iii)分配部分言语(POS)标签,以及(Iv)应用句法注释。但这不适用于非标准数据,因为规则或预先训练的模型尚未适用于所有步骤,并且句法结构的边界是流动的。在标注历史语料库时,前面提到的步骤序列必须按照这个严格的顺序手动完成。因此,我们提供了一个由认知标注过程指导的标注工具,该过程从这条管道后退一步。我们的InterGramm项目显示了对这种工具的需求,我们在该项目中从词法和句法层面研究了中低语德语(MLG)。我们通过使用Cora(Bollmann等人,2014年)开始了我们的注释任务,这是一个成熟的历史数据工具。不幸的是,它不支持句法注释,我们需要追踪13世纪到17世纪不断变化的语法规则。因此,我们扩展了CORA,也捕捉到了不确定性和模糊性(Seemann等人,2017年)。经验表明,在聚合令牌序列和对POS以及构建标签进行比对之前,绑定到令牌级别对人工注释员来说似乎是困难的。自然,语言学家从识别句法模式开始,然后根据相应的语境分配词性标签,并在这个过程中决定一个词位对应多少个标记。此外,他们更喜欢阅读方向的注释,因为这样更容易发现复合词汇或句法结构。因此,我们开发了一个新的具有模式学习支持的标注工具,为注释者提供从以前研究的MLG文本中推断出的建议。
Annotation tools typically use the common text analysis pipeline where (i) tokenization takes place, (ii) End-of-Sentences are detected, (iii) Part-ofSpeech (POS) tags are assigned, and (iv) syntactic annotations are applied. But this does not work for non-standard data where rules or pre-trained models are not yet available for all steps, and boundaries for syntactic constructions are fluid. When annotating historical corpora, the previously mentioned sequence of steps has to be done manually in this strict order. Therefore, we present an annotation tool that is guided by the cognitive annotation process that steps back from this pipeline. The need for such a tool showed up in our project InterGramm where we investigate Middle Low German (MLG) on morphological and syntactic level. We started our annotation task by using CorA (Bollmann et al., 2014), an established tool for historical data. Unfortunately, it does not support syntactic annotations, which we need to trace changing grammar rules from the 13th to 17th century. So we extended CorA, capturing uncertainties and ambiguities as well (Seemann et al., 2017). Experience has shown that being bound to start on token level before aggregating token sequences and aligning POS as well as construction tags appears to be difficult for human annotators. Naturally, linguists start with identifying syntactic patterns, then assign POS tags according to the corresponding context and decide during this process how many tokens belong to one lexeme. Furthermore, they prefer annotating in the direction of reading because it is easier to spot compound lexemes or syntactic constructions. Thus, we developed a new annotation tool with pattern learning support providing the annotators with suggestions inferred from previously studied MLG texts.