Classification and clustering English writing errors based on native language

Classification and clustering English writing errors based on native language
复制标题

基于母语的英语写作错误分类与聚类

DOI:
10.1109/iiai-aai.2014.72
复制
发表时间:
2014
期刊:
Proceedings - 2014 IIAI 3rd International Conference on Advanced Applied Informatics, IIAI-AAI 2014
影响因子:
--
通讯作者:
Hirokawa S.
Hirokawa S.
中科院分区:
--
文献类型:
--
作者:
Flanagan B.;Yin C.;Suzuki T.;Hirokawa S.

文献摘要

相似文献

对于语言学习者来说,确定和反思自己的写作错误是很重要的,这样才能克服缺点。每个语言学习者都有自己独特的写作错误特征,因此有不同的学习需求。本文通过分析外语学习者在语言学习SNS网站Lang-8上的写作错误,探讨母语错误的特征。从Lang-8中收集了142465个句子进行分析。对于每种母语,使用SVM机器学习模型预测的15个错误类别的分数作为每个句子的向量表示。然后对这些分数向量进行聚类,以确定同一句子中的错误共现性。然后对结果进行分析,以确定不同母语的错误特征。
It is important for language learners to determine and reflect on their writing errors in order to overcome weaknesses. Each language learner has their own unique writing error characteristics and therefore has different learning needs. In this paper, we analyze the writing errors of foreign language learners on the language learning SNS website Lang-8 to investigate the characteristics of errors by native language. 142,465 sentences were collected from Lang-8 for analysis. For each native language, the predicted scores of 15 error categories from SVM machine learning models are used as a vector representation of each sentence. These score vectors are then clustered to determine error co-occurrence within the same sentence. The results were then analyzed to determine the error characteristics of different native languages.