What is your Mother Tongue?: Improving Chinese native language identification by cleaning noisy data and adopting BM25

What is your Mother Tongue?: Improving Chinese native language identification by cleaning noisy data and adopting BM25
复制标题

DOI:
10.1109/icbda.2016.7509793
复制
发表时间:
2016-02
期刊:
2016 IEEE International Conference on Big Data Analysis (ICBDA)
影响因子:
--
通讯作者:
Lan Wang;Masahiro Tanaka;H. Yamana
Lan Wang;Masahiro Tanaka;H. Yamana
中科院分区:
其他
文献类型:
--
作者:
Lan Wang;Masahiro Tanaka;H. Yamana

文献摘要

被引文献

相似文献

母语识别(NLI)是从作者的第二语言所写的文章中识别作者母语的过程。在这项工作中,监督模型是建立在一个中国学习者语料库的基础上,以实现这一点。在NLI领域,这是第一项工作:(1)在训练阶段之前自动消除噪声数据,(2)采用BM 25项加权技术对每个特征进行评分。我们还采用了线性支持向量机分类器的层次结构,以实现高精度和国家的最先进的准确率为77.1%,这是大于其他中国NLI方法超过10%。
Native language identification (NLI) is a process by which an author's native language can be identified from essays written in the second language of the author. In this work, a supervised model is built to accomplish this based on a Chinese learner corpus. In the NLI field, this is the first work to (1) eliminate noisy data automatically before the training phase and (2) employ a BM25 term weighting technique to score each feature. We also adopt a hierarchical structure of linear support vector machine classifiers to achieve high accuracy and a state-of-the-art accuracy of 77.1%, which is greater than those of other Chinese NLI methods by over 10%.