Robust, Lexicalized Native Language Identification

Robust, Lexicalized Native Language Identification
复制标题

强大的、词汇化的母语识别

DOI:
--
复制
发表时间:
2012
期刊:
International Conference on Computational Linguistics
影响因子:
--
通讯作者:
Graeme Hirst
Graeme Hirst
中科院分区:
--
文献类型:
--
作者:
Julian Brooke;Graeme Hirst

文献摘要

被引文献

相似文献

以前对母语识别任务的方法(Koppel等人,2005年)仅限于语料库内的小型评估。由于这些都是限制性和不可靠的,我们对任务应用了跨语料库评估。我们展示了词汇特征的有效性,这些特征以前由于语料库内的主题混淆而被避免,并提供了各种选择的详细评估,包括一种简单的偏差适应技术和一些分类器算法。使用一个新的网络语料库作为训练集,我们在一个7种语言的任务上达到了高的分类精度,在两个独立的测试集上表现出了健壮的性能。虽然我们证明了使用交叉验证可以有更高的准确性,但我们提出了强有力的证据,质疑使用标准数据集进行交叉验证评估的有效性。
Previous approaches to the task of native language identification (Koppel et al., 2005) have been limited to small, within-corpus evaluations. Because these are restrictive and unreliable, we apply cross-corpus evaluation to the task. We demonstrate the efficacy of lexical features, which had previously been avoided due to the within-corpus topic confounds, and provide a detailed evaluation of various options, including a simple bias adaptation technique and a number of classifier algorithms. Using a new web corpus as a training set, we reach high classification accuracy for a 7-language task, performance which is robust across two independent test sets. Although we show that even higher accuracy is possible using crossvalidation, we present strong evidence calling into question the validity of cross-validation evaluation using the standard dataset.