Automatic Linguistic Annotation ofLarge Scale L2 Databases: The EF-Cambridge Open Language Database(EFCamDat)

Automatic Linguistic Annotation ofLarge Scale L2 Databases: The EF-Cambridge Open Language Database(EFCamDat)
复制标题

大规模 L2 数据库的自动语言注释:EF-剑桥开放语言数据库 (EFCamDat)

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
A. Korhonen
A. Korhonen
中科院分区:
--
文献类型:
--
作者:
J. Geertzen;T. Alexopoulou;A. Korhonen

文献摘要

被引文献

相似文献

自然主义的学习者产出是二语习得研究的重要经验资源。一些开创性的工作已经产生了支持二语习得研究的有价值的第二语言(L2)资源。1这些资源的一个共同的局限性是缺乏跨熟练度谱的具有不同背景的众多说话者的个体纵向数据,这对于理解纵向发展中个体差异的本质至关重要。2第二个限制是用语言注释的数据数量相对有限。信息(例如,词汇、形态句法、语义特征等)以支持二语习得假说的研究,并获得不同语言现象的发展模式。在有注释的情况下,注释往往是手工获得的,这种情况直接限制了在合理的人力资源和合理的时间内可以注释的数据数量。自然语言处理(NLP)工具可以为词性(POS)和句法结构提供自动注释,并且确实越来越多地应用于各种上下文中的学习者语言。计算机辅助语言学习(CALL)系统使用解析器和其他NLP工具自动检测学习者错误并提供相应的反馈。3一些工作旨在调整解析工具提供的注释以准确描述学习者语法(Dickinson & Lee,2009)或评估解析器对学习者语言的性能以及学习者错误对解析器的影响。Krivanek和Meurers(2011)比较了两种解析方法,一种使用手工制作的词典,另一种在语料库上训练。他们发现,前者在恢复主要的语法依赖关系方面更成功,而后者在恢复可选的附加关系方面更成功。Ott和Ziai(2010)评估了在母语德语上训练的依赖解析器的性能(MaltParser; Nivre等人,2007年)对106名学习者回答L2德语理解任务。他们的研究表明,虽然一些错误可能对解析器造成问题(例如,省略限定动词)许多其它的(例如,错误的词序)可以被鲁棒地解析,从而导致整体高性能分数。在本文中,我们有两个目标。首先,我们介绍了一个新的英语L2数据库,EF剑桥开放语言数据库,以下简称EFCAMDAT。EFCAMDAT由剑桥大学理论与应用语言学系与国际教育组织EF Education First合作开发。它包含了提交给英国城的作品,
∗Naturalistic learner productions are an important empirical resource for SLA research. Some pioneering works have produced valuable second language (L2) resources supporting SLA research.1 One common limitation of these resources is the absence of individual longitudinal data for numerous speakers with different backgrounds across the proficiency spectrum, which is vital for understanding the nature of individual variation in longitudinal development.2 A second limitation is the relatively restricted amounts of data annotated with linguistic information (e.g., lexical, morphosyntactic, semantic features, etc.) to support investigation of SLA hypotheses and obtain patterns of development for different linguistic phenomena. Where available, annotations tend to be manually obtained, a situation posing immediate limitations to the quantity of data that could be annotated with reasonable human resources and within reasonable time. Natural Language Processing (NLP) tools can provide automatic annotations for parts-of-speech (POS) and syntactic structure and are indeed increasingly applied to learner language in various contexts. Systems in computer-assisted language learning (CALL) have used a parser and other NLP tools to automatically detect learner errors and provide feedback accordingly.3 Some work aimed at adapting annotations provided by parsing tools to accurately describe learner syntax (Dickinson & Lee, 2009) or evaluated parser performance on learner language and the effect of learner errors on the parser. Krivanek and Meurers (2011) compared two parsing methods, one using a hand-crafted lexicon and one trained on a corpus. They found that the former is more successful in recovering the main grammatical dependency relations whereas the latter is more successful in recovering optional, adjunction relations. Ott and Ziai (2010) evaluated the performance of a dependency parser trained on native German (MaltParser; Nivre et al., 2007) on 106 learner answers to a comprehension task in L2 German. Their study indicates that while some errors can be problematic for the parser (e.g., omission of finite verbs) many others (e.g., wrong word order) can be parsed robustly, resulting in overall high performance scores. In this paper we have two goals. First, we introduce a new English L2 database, the EF Cambridge Open Language Database, henceforth EFCAMDAT. EFCAMDAT was developed by the Department of Theoretical and Applied Linguistics at the University of Cambridge in collaboration with EF Education First, an international educational organization. It contains writings submitted to Englishtown, the