What is your Mother Tongue?: Improving Chinese native language identification by cleaning noisy data and adopting BM25
What is your Mother Tongue?: Improving Chinese native language identification by cleaning noisy data and adopting BM25
复制标题
DOI:
10.1109/icbda.2016.7509793
复制
发表时间:
2016-02
期刊:
影响因子:
--
通讯作者:
Lan Wang;Masahiro Tanaka;H. Yamana
中科院分区:
文献类型:
--
作者:
Lan Wang;Masahiro Tanaka;H. Yamana
Native language identification (NLI) is a process by which an author's native language can be identified from essays written in the second language of the author. In this work, a supervised model is built to accomplish this based on a Chinese learner corpus. In the NLI field, this is the first work to (1) eliminate noisy data automatically before the training phase and (2) employ a BM25 term weighting technique to score each feature. We also adopt a hierarchical structure of linear support vector machine classifiers to achieve high accuracy and a state-of-the-art accuracy of 77.1%, which is greater than those of other Chinese NLI methods by over 10%.