Annotating an Arabic Learner Corpus for Error

Annotating an Arabic Learner Corpus for Error
复制标题

注释阿拉伯语学习者语料库中的错误

DOI:
--
复制
发表时间:
2008
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Eileen Fitzpatrick
Eileen Fitzpatrick
中科院分区:
--
文献类型:
--
作者:
Ghazi Abuhakema;Reem Faraj;Anna Feldman;Eileen Fitzpatrick

文献摘要

被引文献

相似文献

本文描述了一个正在进行的项目,在这个项目中,我们收集了一个阿拉伯语学习者语料库,开发了一个错误标注标签集,并对数据进行了计算机辅助错误分析(CEA)。我们采用了法语中介语数据库Frida标记集(Granger,2003a)来适应这些数据。我们选择Frida是为了遵循一个已知的标准,并看看从法语标记集转移到阿拉伯语标记集所需的更改是否会让我们衡量两种语言在学习者难度方面的距离。目前的课本数量不断增加,其中包括中级和高级学生的作文。我们描述了对这样的语料库的需求,我们收集的学习者数据和我们开发的标签集。我们还描述了熟练程度和正在进行的工作的错误频率分布。
This paper describes an ongoing project in which we are collecting a learner corpus of Arabic, developing a tagset for error annotation and performing Computer-aided Error Analysis (CEA) on the data. We adapted the French Interlanguage Database FRIDA tagset (Granger, 2003a) to the data. We chose FRIDA in order to follow a known standard and to see whether the changes needed to move from a French to an Arabic tagset would give us a measure of the distance between the two languages with respect to learner difficulty. The current collection of texts, which is constantly growing, contains intermediate and advanced-level student writings. We describe the need for such corpora, the learner data we have collected and the tagset we have developed. We also describe the error frequency distribution of both proficiency levels and the ongoing work.