More efficient and accurate automatic speech recognition
More efficient and accurate automatic speech recognition
批准号:
RGPIN-2018-05226
负责人:
OShaughnessy, Douglas
金额:
$2.04万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Current Automatic Speech Recognition (ASR) uses stochastic methods that exclude much of what is known about human speech production and perception. For 30 years, ASR has used Hidden Markov Models (HMMs) and now Deep Neural Networks (DNNs). Both are engineering approaches emphasizing recognition accuracy, but tolerating ever increasing cost (computer memory and processing). NNs have existed for decades, but applications were mostly limited to 3-level multilayer perceptrons, with limited capacity to handle the wide range of variability in speech (sources, channels, speakers, contexts, environments). Despite much increased use of ASR (e.g, Siri, Alexa), performance is still not near human levels, especially for noisy conditions (e.g., many cases where prior model training is limited). Continuing recent DNN ASR research is unlikely to approach acceptable accuracy in many cases unless major changes are made to the methodology.Early ASR methodology in the 1970s used mostly expert-system (ES) approaches, exploiting ideas of how vocal-tract resonances (called formants) related to the phonemes intended by speakers, and focused on the spectral peaks of speech, as this is how the ear filters speech inside the cochlea. In the early 1980s, HMMs took over the ASR field as they were much better at handling variability than simple “if-then” algorithms. Nonetheless, if one could track significant aspects of resonances reliably in poor acoustic conditions (that human listeners handle well), then useful ASR decisions could be made far at lower cost than with recent end-to-end DNN approaches. It is here proposed to combine structural and stochastic information in ASR, exploiting well-known (but, in ASR, little used) knowledge of how humans do speech communication.Another major deficiency of ASR is its lack of use of intonation, despite all evidence that such facilitates human speech communication. Human intonation in speech production (which is clearly exploited by human listeners) has long time ranges, making such information very difficult to track in the current systems that rely on either raw speech (at 8000 samples/s or higher) or 10-ms frames of spectral data. The relative success of modern HMM and DNN approaches show that one can succeed (to a certain level of performance) without using intonation, as many practical speech inputs to ASR are simple and short phrases (and in good quality environments). Nonetheless, proper use of intonation in ASR would surely raise recognition accuracy, just as including language models (LM) into ASR in the 1980s did. We will improve robustness in ASR to common acoustic degradations, have greater efficiency, and exploit intonation. The long-term objective: accurate and efficient ASR, approaching that of human listeners. Short-term objectives: 1) a better spectral measure than filter-bank energies, 2) faster and better adaptation, 3) integrate aspects of intonation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
More efficient and accurate automatic speech recognition
-
批准号:RGPIN-2018-05226
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2021
-
负责人:OShaughnessy, Douglas
-
依托单位:
More efficient and accurate automatic speech recognition
-
批准号:RGPIN-2018-05226
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2020
-
负责人:OShaughnessy, Douglas
-
依托单位:
More efficient and accurate automatic speech recognition
-
批准号:RGPIN-2018-05226
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2019
-
负责人:OShaughnessy, Douglas
-
依托单位:
More efficient and accurate automatic speech recognition
-
批准号:RGPIN-2018-05226
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2018
-
负责人:OShaughnessy, Douglas
-
依托单位:
More Accurate and Efficient Analysis for Automatic Speech Recognition
-
批准号:914-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2017
-
负责人:OShaughnessy, Douglas
-
依托单位:
More Accurate and Efficient Analysis for Automatic Speech Recognition
-
批准号:914-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2016
-
负责人:OShaughnessy, Douglas
-
依托单位:
More Accurate and Efficient Analysis for Automatic Speech Recognition
-
批准号:914-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2015
-
负责人:OShaughnessy, Douglas
-
依托单位:
More Accurate and Efficient Analysis for Automatic Speech Recognition
-
批准号:914-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2014
-
负责人:OShaughnessy, Douglas
-
依托单位:
More Accurate and Efficient Analysis for Automatic Speech Recognition
-
批准号:914-2013
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.11万
-
财政年份:2013
-
负责人:OShaughnessy, Douglas
-
依托单位:
Improving basic methods of automatic speech recognition
-
批准号:914-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.83万
-
财政年份:2012
-
负责人:OShaughnessy, Douglas
-
依托单位:
Improving basic methods of automatic speech recognition
-
批准号:914-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.83万
-
财政年份:2011
-
负责人:OShaughnessy, Douglas
-
依托单位:
Improving basic methods of automatic speech recognition
-
批准号:914-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.83万
-
财政年份:2010
-
负责人:OShaughnessy, Douglas
-
依托单位:
Improving basic methods of automatic speech recognition
-
批准号:914-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.83万
-
财政年份:2009
-
负责人:OShaughnessy, Douglas
-
依托单位:
Improving basic methods of automatic speech recognition
-
批准号:914-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.83万
-
财政年份:2008
-
负责人:OShaughnessy, Douglas
-
依托单位:
Automatic speech recognition for mobile and personal applications
-
批准号:336627-2006
-
项目类别:Strategic Projects - Group
-
资助金额:$8.74万
-
财政年份:2008
-
负责人:OShaughnessy, Douglas
-
依托单位:
Improving automatic speech recognition techniques
-
批准号:914-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.92万
-
财政年份:2007
-
负责人:OShaughnessy, Douglas
-
依托单位:
Automatic speech recognition for mobile and personal applications
-
批准号:336627-2006
-
项目类别:Strategic Projects - Group
-
资助金额:$8.74万
-
财政年份:2007
-
负责人:OShaughnessy, Douglas
-
依托单位:
Automatic speech recognition for mobile and personal applications
-
批准号:336627-2006
-
项目类别:Strategic Projects - Group
-
资助金额:$8.74万
-
财政年份:2006
-
负责人:OShaughnessy, Douglas
-
依托单位:
Improving automatic speech recognition techniques
-
批准号:914-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.92万
-
财政年份:2006
-
负责人:OShaughnessy, Douglas
-
依托单位:
Improving automatic speech recognition techniques
-
批准号:914-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.92万
-
财政年份:2005
-
负责人:OShaughnessy, Douglas
-
依托单位:
国内基金
海外基金
固定参数可解算法在平面图问题的应用以及和整数线性规划的关系
-
批准号:60973026
-
项目类别:面上项目
-
资助金额:32.0万元
-
批准年份:2009
-
负责人:鲁道夫
-
依托单位: