How Do Humans Process and Recognize Speech?

How Do Humans Process and Recognize Speech?
复制标题

DOI:
10.1109/89.326615
复制
发表时间:
1994-10-01
期刊:
IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING
影响因子:
--
通讯作者:
Allen, Jont B.
Allen, Jont B.
中科院分区:
其他
文献类型:
--
作者:
Allen, Jont B.

文献摘要

被引文献

相似文献

在自动语音识别(ASR)硬件的性能在准确性和鲁棒性方面超过人类性能之前,我们将通过理解人类语音识别(HSR)背后的基本原理来获益。1918年至1950年,哈维·弗莱彻和他的同事们在贝尔实验室对这个问题进行了详尽的研究。这些研究的动机是量化电话厂的语音质量,以提高语音清晰度和偏好。为此,他和他的团队研究了过滤和噪声对无意义辅音-元音-辅音(CVC)音节、单词和句子的语音识别准确性的影响。弗莱彻用“清晰度”来表示正确识别无意义声音的概率,用“可理解性”来表示正确识别单词(声音有意义)的概率。1919年,弗莱彻发现了一种方法,可以将过滤语音的清晰度数据转换为加性密度函数D(f),并找到了一个公式,可以准确地预测平均清晰度。D(f)下的区域称为“清晰度指数”。弗莱彻接着发现了无意义语音、单词和句子的识别错误之间的关系。Boothroyd和Bronkhorst等人最近对这项工作进行了回顾和部分复制。总的来说,这些研究告诉我们很多关于人类如何处理和识别语音的信息。
Until the performance of automatic speech recognition (ASR) hardware surpasses human performance in accuracy and robustness, we stand to gain by understanding the basic principles behind human speech recognition (HSR). This problem was studied exhaustively at Bell Labs between the years of 1918 and 1950 by Harvey Fletcher and his colleagues. The motivation for these studies was to quantify the quality of speech sounds in the telephone plant to both improve speech intelligibility and preference. To do this he and his group studied the effects of filtering and noise on speech recognition accuracy for nonsense consonant-vowel-consonant (CVC) syllables, words, and sentences. Fletcher used the term "articulation" as the probability of correct recognition for nonsense sounds, and "intelligibility" as the probability of correction recognition for words (sounds having meaning). In 1919, Fletcher found a way to transform articulation data for filtered speech into an additive density function D(f) and found a formula that accurately predicts the average articulation. The area under D(f) is called the "articulation index." Fletcher then went on to find relationships between the recognition errors for the nonsense speech sounds, words, and sentences. This work has recently been reviewed and partially replicated by Boothroyd and by Bronkhorst, et al. Taken as a whole, these studies tell us a great deal about how humans process and recognize speech sounds.