Objective Intelligibility Assessment by Automated Segmental and Suprasegmental Listening Error Analysis.

Objective Intelligibility Assessment by Automated Segmental and Suprasegmental Listening Error Analysis.
复制标题

通过自动分段和超分段听力错误分析进行客观清晰度评估。

DOI:
10.1044/2019_jslhr-s-19-0119
复制
发表时间:
2019
期刊:
Journal of speech, language, and hearing research : JSLHR
影响因子:
--
通讯作者:
Liss,Julie
Liss,Julie
中科院分区:
--
文献类型:
--
作者:
Jiao,Yishan;LaCross,Amy;Berisha,Visar;Liss,Julie

文献摘要

被引文献

相似文献

目的主观的语音清晰度评估通常比依赖于成绩单评分的更客观的方法更受欢迎。这在一定程度上是因为从转录的语音中提取客观指标需要大量的体力劳动。在这项研究中,我们提出了一种自动评分转录本的方法,该方法提供了源于片段和超片段贡献的可理解性退化的整体和客观表示,并且与人类感知相对应。方法819名听者通过Mechanical Turk对73名构音障碍说话者产生的短语进行正字法转录,得到63840个短语转录结果。开发了一个协议来过滤转录本,然后使用用于测量音素和词汇分割错误的新算法自动分析转录本。将结果与随机选择的40个转录短语样本集上的手动标签进行比较,以评估有效性。进行了线性回归分析,以检验自动化指标预测严重程度和单词准确性的感知评级的效果。结果在样本集上,自动度量方法对音位错误的度量与人工标注的相关系数达到0.90,对词法分词错误的识别和编码准确率达到100%。线性回归模型发现,估计的指标可以预测感知严重程度和单词准确性的显著部分方差。研究结果表明,客观的语音可理解性评估方法有望在多个分析层次上识别可理解性退化。
PurposeSubjective speech intelligibility assessment is often preferred over more objective approaches that rely on transcript scoring. This is, in part, because of the intensive manual labor associated with extracting objective metrics from transcribed speech. In this study, we propose an automated approach for scoring transcripts that provides a holistic and objective representation of intelligibility degradation stemming from both segmental and suprasegmental contributions, and that corresponds with human perception.MethodPhrases produced by 73 speakers with dysarthria were orthographically transcribed by 819 listeners via Mechanical Turk, resulting in 63,840 phrase transcriptions. A protocol was developed to filter the transcripts, which were then automatically analyzed using novel algorithms developed for measuring phoneme and lexical segmentation errors. The results were compared with manual labels on a randomly selected sample set of 40 transcribed phrases to assess validity. A linear regression analysis was conducted to examine how well the automated metrics predict a perceptual rating of severity and word accuracy.ResultsOn the sample set, the automated metrics achieved 0.90 correlation coefficients with manual labels on measuring phoneme errors, and 100% accuracy on identifying and coding lexical segmentation errors. Linear regression models found that the estimated metrics could predict a significant portion of the variance in perceptual severity and word accuracy.ConclusionsThe results show the promising development of an objective speech intelligibility assessment that identifies intelligibility degradation on multiple levels of analysis.