Is text-to-speech synthesis ready for use in computer-assisted language learning?

Is text-to-speech synthesis ready for use in computer-assisted language learning?
复制标题

DOI:
10.1016/j.specom.2008.12.004
复制
发表时间:
2009-10-01
影响因子:
3.2
通讯作者:
Handley, Zoee
Handley, Zoee
中科院分区:
计算机科学3区
文献类型:
--
作者:
Handley, Zoee

文献摘要

被引文献

相似文献

文本到语音(TTS)合成,即从文本输入生成语音,为计算机辅助语言学习(CALL)环境中的学习者提供了另一种提供口语输入的方法。事实上,许多潜在的好处(语音模型的创建和编辑、语音模型的生成和按需反馈等)。和使用(对话词典、对话文本、听写、发音训练、对话伙伴等)提出了在CALL中实现TTS合成的方法。然而,在Call中使用TTS合成并没有被广泛接受,只有少数几个应用程序进入了市场。一个潜在的原因是,TTS合成还没有得到充分的评估。以前对用于CALL的TTS合成的评估只涉及TTS合成的可理解性。然而,CALL对TTS合成输出的可理解性、自然性、准确性、注册性和表现力提出了要求。在本文中,针对TTS合成系统在呼叫应用中可能承担的三种不同角色,即(1)阅读机、(2)发音模型和(3)会话伙伴[Handley,Z.,Hamel,M.-J.,2005],对四个最先进的法语TTS合成系统的输出质量的上述方面进行了评估。建立用于计算机辅助语言学习(CALL)的语音合成基准方法。《语言学习与技术》期刊9(3),99-119。检索自:http://llt.msu.edu/vol9num3/handley/default.html.].这一评估的结果表明,最好的TTS合成系统已经准备好用于它们可以‘增值’的应用中,即利用TTS合成的独特能力来按需生成语音模型。对话伙伴就是这种应用程序的一个例子。为了完全满足CALL的要求,需要进一步注意准确性和自然性,特别是在韵律层面和表现力上。(C)2008爱思唯尔B.V.保留所有权利。
Text-to-speech (TTS) synthesis, the generation of speech from text input, offers another means of providing spoken language input to learners in Computer-Assisted Language Learning (CALL) environments. Indeed, many potential benefits (ease of creation and editing of speech models, generation of speech models and feedback on demand, etc.) and uses (talking dictionaries, talking texts, dictation, pronunciation training, dialogue partner, etc.) of TTS synthesis in CALL have been put forward. Yet, the use of TTS synthesis in CALL is not widely accepted and only a few applications have found their way onto the market. One potential reason for this is that TTS synthesis has not been adequately evaluated for this purpose. Previous evaluations of TTS synthesis for use in CALL, have only addressed the comprehensibility of TTS synthesis. Yet, CALL places demands on the comprehensibility, naturalness, accuracy, register and expressiveness of the output of TTS synthesis. In this paper, the aforementioned aspects of the quality of the output of four state-of-the-art French TTS synthesis systems are evaluated with respect to their use in the three different roles that TTS synthesis systems may assume within CALL applications, namely: (1) reading machine, (2) pronunciation model and (3) conversational partner [Handley, Z., Hamel, M.-J., 2005. Establishing a methodology for benchmarking speech synthesis for computer-assisted language learning (CALL). Language Learning and Technology Journal 9(3), 99-119. Retrieved from: http://llt.msu.edu/vol9num3/handley/default.html.]. The results of this evaluation suggest that the best TTS synthesis systems are ready for use in applications in which they 'add value' to CALL, i.e. exploit the unique capacity of TTS synthesis to generate speech models on demand. An example of such an application is a dialogue partner. In order to fully meet the requirements of CALL, further attention needs to be paid to accuracy and naturalness, in particular at the prosodic level, and expressiveness. (C) 2008 Elsevier B.V. All rights reserved.