Comparison of Static and Time-Sequential Features in Automatic Fluency Detection of Spontaneous Speech

Comparison of Static and Time-Sequential Features in Automatic Fluency Detection of Spontaneous Speech
复制标题

自发语音自动流利度检测中静态特征与时序特征的比较

DOI:
10.1109/o-cocosda202152914.2021.9660601
复制
发表时间:
2021
期刊:
Proceedings of the 24th Conference of the Oriental COCOSDA
影响因子:
--
通讯作者:
Nishizaki Hiromitsu
Nishizaki Hiromitsu
中科院分区:
--
文献类型:
--
作者:
Deng Huaijin;Utsuro Takehito;Kobayashi Akio;Nishizaki Hiromitsu

文献摘要

参考文献

相似文献

在不流利检测方面已有大量的研究。然而,他们中的大多数都集中在词汇线索上,很少强调不同的声学特征如何有助于提高表现。我们描述了一个自动流畅性评价框架的自发语音。我们不仅研究了从转录中提取的词汇特征,还考虑了音频数据中的时间序列和静态声学特征,包括能量和语音相关特征。本研究试图揭示不同的声学特征是如何影响语音流畅性评价的。使用lstm和DNN架构对自发日语语料库进行了评估。评价结果表明,在检测流利语音时,将四个时间序列声学特征与静态词汇特征相结合的方法效果最好。另一方面,在检测不流畅语音时,静态抖动/闪烁特征有助于在相对较高的下界提高精度。
There have been lots of previous studies on disfluency detection. However, most of them focus on lexical cues, and little emphasis is placed on how diverse acoustic features contribute to improving the performance. We describe a framework for automatic fluency evaluation of spontaneous speech. We investigate not only lexical features extracted from transcription, but also consider time-sequential and static acoustic features from audio data, including energy and voicing-related features. This work tries to reveal how diverse acoustic features contribute to the performance of speech fluency evaluation. The proposed framework was evaluated with the Corpus of Spontaneous Japanese using LSTMs and DNN architectures. Evaluation results showed that when detecting fluent speech, combining four time-sequential acoustic features with static lexical features achieved the best performance. When detecting disfluent speech, on the other hand, static jitter/shimmer features helped to improve the precision at relatively high lower bounds.
DOI: 10.1109/tasl.2006.878255
发表时间: 2006-09
期刊: IEEE Transactions on Audio, Speech, and Language Processing
影响因子: --
作者:
Yang Liu;Elizabeth Shriberg;A. Stolcke;D. Hillard;Mari Ostendorf;M. Harper
通讯作者: Yang Liu;Elizabeth Shriberg;A. Stolcke;D. Hillard;Mari Ostendorf;M. Harper
语音识别器对于课堂演讲语音的特征分析有用吗?
DOI: --
发表时间: 2008
期刊: the Proc. of the 9th annual conference of the International Speech Communication Association (INTERSPEECH2008)
影响因子: --
作者:
Kenji Kobayashi;Mitsuhiro Somiya;Hiromitsu Nishizaki and Yoshihiro Sekiguchi
通讯作者: Hiromitsu Nishizaki and Yoshihiro Sekiguchi
DOI: 10.21437/interspeech.2009-103
发表时间: 2009
期刊: --
影响因子: --
作者:
Björn Schuller;S. Steidl;A. Batliner
通讯作者: Björn Schuller;S. Steidl;A. Batliner
英语口语流利度自动评估
DOI: 10.1109/icassp.2009.4960712
发表时间: 2009
期刊: 2009 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子: --
作者:
Om Deshmukh;Kundan Kandhway;Ashish Verma;Kartik Audhkhasi
通讯作者: Kartik Audhkhasi