Automated assessment of second language comprehensibility: Review, training, validation, and generalization studies

Automated assessment of second language comprehensibility: Review, training, validation, and generalization studies
复制标题

DOI:
10.1017/s0272263122000080
复制
发表时间:
2022-03
影响因子:
4.1
通讯作者:
Kazuya Saito;Konstantinos Macmillan;Magdalena Kachlicka;Takuya Kunihara;N. Minematsu
Kazuya Saito;Konstantinos Macmillan;Magdalena Kachlicka;Takuya Kunihara;N. Minematsu
中科院分区:
人文科学1区
文献类型:
--
作者:
Kazuya Saito;Konstantinos Macmillan;Magdalena Kachlicka;Takuya Kunihara;N. Minematsu

文献摘要

被引文献

相似文献

尽管许多学者强调可理解性作为第二语言语音训练、测试和发展的生态有效目标的相对重要性,但引发听者的判断是耗时的。随着应用语言学中对更有效的L2语音评级方法的研究的呼吁,以及对语音工程中自发脱稿语音使用机器学习的日益关注,本研究探讨了建立快速可靠的自动可理解性评估的可能性。回归模型展示了一组语音(L1和L2语音之间的最大后验概率和间隙),韵律(音高和强度变化)和时间测量(清晰度,停顿频率),显着预测如何天真的听众直观地判断低,中,高,和nativelike理解100 L1和L2扬声器的图片描述。相关性的强度(机器与人类评分的r = 0.823)与天真听众的评分者间一致性(人类与人类的r = 0.760)相当。当该模型应用于45个L1和L2说话者的新数据集(r = .827)并在更自由构建的面试任务条件下进行测试(r = .809)时,结果被成功复制。
Abstract Whereas many scholars have emphasized the relative importance of comprehensibility as an ecologically valid goal for L2 speech training, testing, and development, eliciting listeners’ judgments is time-consuming. Following calls for research on more efficient L2 speech rating methods in applied linguistics, and growing attention toward using machine learning on spontaneous unscripted speech in speech engineering, the current study examined the possibility of establishing quick and reliable automated comprehensibility assessments. Orchestrating a set of phonological (maximum posterior probabilities and gaps between L1 and L2 speech), prosodic (pitch and intensity variation), and temporal measures (articulation rate, pause frequency), the regression model significantly predicted how naïve listeners intuitively judged low, mid, high, and nativelike comprehensibility among 100 L1 and L2 speakers’ picture descriptions. The strength of the correlation (r = .823 for machine vs. human ratings) was comparable to naïve listeners’ interrater agreement (r = .760 for humans vs. humans). The findings were successfully replicated when the model was applied to a new dataset of 45 L1 and L2 speakers (r = .827) and tested under a more freely constructed interview task condition (r = .809).