Fifty years of progress in speech and speaker recognition

Fifty years of progress in speech and speaker recognition
复制标题

语音和说话人识别领域五十年的进展

DOI:
10.1121/1.4784967
复制
发表时间:
2004
影响因子:
2.4
通讯作者:
S. Furui
S. Furui
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
S. Furui

文献摘要

参考文献

被引文献

相似文献

语音和说话人识别技术在过去的50年里取得了非常重大的进展。这些进展可以概括为以下变化:(1)从模板匹配到基于语料库的统计建模,例如,HMM和n-gram,(2)从滤波器组/谱共振到倒谱特征(Cepstrum + DCepstrum + DDCepstrum),(3)从启发式时间归一化到DTW/DP匹配,(4)从基于gdistanceh的方法到基于似然的方法,(5)从最大似然到判别方法,例如,MCE/GPD和MMI,(6)从孤立词到连续语音识别,(7)从小词汇到大词汇识别,(8)从上下文无关单元到上下文相关单元识别,(9)从干净语音到嘈杂/电话语音识别,(10)从单个说话人到说话人无关/自适应识别,(11)从独白到对话/会话识别,(12)从读语音到自发语音识别,(13)从识别到理解,(14)从单模态(仅音频信号)到多模态(音频/视频)语音识别,(15)从硬件识别器到软件识别器,(16)从没有商业应用到许多实际的商业应用。这些进展大多发生在语音识别和说话人识别两个领域。大多数技术变化都是为了提高识别的鲁棒性,包括上面没有提到的许多其他重要技术。
Speech and speaker recognition technology has made very significant progress in the past 50 years. The progress can be summarized by the following changes: (1) from template matching to corpus‐base statistical modeling, e.g., HMM and n‐grams, (2) from filter bank/spectral resonance to Cepstral features (Cepstrum + DCepstrum + DDCepstrum), (3) from heuristic time‐normalization to DTW/DP matching, (4) from gdistanceh‐based to likelihood‐based methods, (5) from maximum likelihood to discriminative approach, e.g., MCE/GPD and MMI, (6) from isolated word to continuous speech recognition, (7) from small vocabulary to large vocabulary recognition, (8) from context‐independent units to context‐dependent units for recognition, (9) from clean speech to noisy/telephone speech recognition, (10) from single speaker to speaker‐independent/adaptive recognition, (11) from monologue to dialogue/conversation recognition, (12) from read speech to spontaneous speech recognition, (13) from recognition to understanding, (14) from single‐modality (audio signal only) to multi‐modal (audio/visual) speech recognition, (15) from hardware recognizer to software recognizer, and (16) from no commercial application to many practical commercial applications. Most of these advances have taken place in both the fields of speech recognition and speaker recognition. The majority of technological changes have been directed toward the purpose of increasing robustness of recognition, including many other additional important techniques not noted above.
通过并行 Potts 分割进行神经元识别。
DOI: 10.1073/pnas.0230490100
发表时间: 2003
期刊: Proceedings of the National Academy of Sciences of the United States of America.
影响因子: --
作者:
Peng,S;Urbanc,B;Cruz,L;Hyman,BT;Stanley,HE
通讯作者: Stanley,HE