Open answer scoring for S-CAT automated speaking test system using support vector regression

Open answer scoring for S-CAT automated speaking test system using support vector regression
复制标题

DOI:
--
复制
发表时间:
2012-12
期刊:
Proceedings of The 2012 Asia Pacific Signal and Information Processing Association Annual Summit and Conference
影响因子:
--
通讯作者:
Yutaka Ono;Misuzu Otake;T. Shinozaki;R. Nisimura;Takeshi Yamada;K. Ishizuka;Y. Horiuchi;S. Kuroiwa;S. Imai
Yutaka Ono;Misuzu Otake;T. Shinozaki;R. Nisimura;Takeshi Yamada;K. Ishizuka;Y. Horiuchi;S. Kuroiwa;S. Imai
中科院分区:
其他
文献类型:
--
作者:
Yutaka Ono;Misuzu Otake;T. Shinozaki;R. Nisimura;Takeshi Yamada;K. Ishizuka;Y. Horiuchi;S. Kuroiwa;S. Imai

文献摘要

相似文献

我们正在开发S-CAT计算机考试系统,这将是第一个自动自适应日语口语考试。考生的口语能力评分采用语音处理技术,没有人工评分。使用计算机进行评分,可以大大降低评分成本,为语言学习者评估自己的学习状况提供了一种方便的手段。虽然S-CAT考试有几类问题,但开放性问题在技术上是最具挑战性的,因为考生可以自由地谈论给定的话题或就给定的材料进行辩论。针对这一问题,我们提出使用具有多种特征的支持向量回归(SVR)。一些特征依赖于语音识别假设,而另一些则不是。支持向量回归比多元回归具有更强的鲁棒性,当使用390维特征组合所有特征时,得到的结果最好。在流畅性、准确性、内容和丰富度方面,人类评分与SVR估计得分之间的相关系数分别为0.878、0.847、0.853和0.872。
We are developing S-CAT computer test system that will be the first automated adaptive speaking test for Japanese. The speaking ability of examinees is scored using speech processing techniques without human raters. By using computers for the scoring, it is possible to largely reduce the scoring cost and provide a convenient means for language learners to evaluate their learning status. While the S-CAT test has several categories of question items, open answer question is technically the most challenging one since examinees freely talk about a given topic or argue something for a given material. For this problem, we proposed to use support vector regression (SVR) with various features. Some of the features rely on speech recognition hypothesis and others do not. SVR is more robust than multiple regression and the best result was obtained when 390 dimensional features that combine everything were used. The correlation coefficients between human rated and SVR estimated scores were 0.878, 0.847, 0.853, and 0.872 for fluency, accuracy, content, and richness measures, respectively.