A method for reducing burden imposed on human raters in the construction of automated scoring systems for second language learners’ speech

A method for reducing burden imposed on human raters in the construction of automated scoring systems for second language learners’ speech
复制标题

一种在构建第二语言学习者语音自动评分系统时减轻人工评分者负担的方法

DOI:
--
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Yusuke Kondo
Yusuke Kondo
中科院分区:
--
文献类型:
--
作者:
Yusuke Kondo

文献摘要

相似文献

人们已经尝试构建第二语言学习者语音的自动评分系统。在这些系统构建的初始阶段,研究了人类评分者的分数与计算机可测量的语音特征之间的关系,以获得预测公式:一旦我们获得公式,就可以使用语音特征来预测考生的分数。虽然提出了计算机化评估作为减轻评分员负担的解决方案之一,但在系统构建的初始阶段,需要大量学习者的语音数据和人工评分员给出的分数。评分者需要评估大量的语音样本。为了解决这一问题,本研究提出了一种方法,通过少量的语音数据和人工评分来预测大量未评分的语音数据的得分。本研究使用的演讲数据是101个亚洲英语学习者的朗读演讲。利用随机选取的5个演讲的2个语音特征,基于Expectation-Maximum (EM)算法预测剩余86个演讲的得分。人类评分者给出的分数与算法预测的分数之间存在适度的相关性(约为0.60)。
Attempts have been made to construct automated scoring systems for second language learners’ speech. In the initial stage of the construction of these systems, the relationship is investigated between the scores by human raters and speech characteristics that are measurable by computer in order to obtain prediction formulae: Once we obtain the formulae, examinees’ scores can be predicted using speech characteristics. Although the computerized assessment is proposed as one of the solutions to reduce the raters’ burden, the initial stage of the system construction requires a large amount of learners’ speech data with the scores given by human raters. The raters need to evaluate a large set of speech samples. To solve this problem, this study proposes a method for predicting the scores of a large set of unscored speech data by a small set of speech data with the human rating. The speech data used in this study are 101 read-aloud speeches given by Asian learners of English. Using two speech characteristics of five speeches randomly selected, the scores of the remaining 86 speeches are predicted, based on Expectation-Maximum (EM) algorithm. The moderate correlation was found between the scores given by the human raters and the ones predicted by the algorithm (around .60).