Deep structured speech models
Deep structured speech models
批准号:
RGPIN-2021-02652
负责人:
Boulianne, Gilles
金额:
$2.77万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
State-of-the-art speech recognition models are either end-to-end models or hybrid models. End-to-end models are entirely based on deep neural networks (DNN). Hybrid models combine an underlying structure of weighted finite-state automata (WFSA) with a surface layer of deep neural networks. When trained on large enough amounts of annotated recordings similar to the test data, end-to-end models can outperform hybrid models. For mismatched or smaller training data, hybrid models are often a better choice since they can use prior linguistic knowledge encoded in their structure to better generalize and avoid overfitting. However, in hybrid models, finite-state automata and deep neural components are not integrated; they are trained independently with different objectives and algorithms. As a result, a performance gap remains even when enough training data is available. In practice, hybrid models are complicated to train, require specialized coding, and cannot be easily integrated with common deep learning frameworks such as Pytorch or Tensorflow. Thus their implementations lag behind the latest developments in general deep learning. Recent proposals for differentiable automata suggest how they could be integrated with deep neural networks into a single model with an end-to-end differentiable loss, in a way that is efficient in time and space to be scalable enough for speech problems. This opens up several areas of research which are yet almost unexplored for speech modelling. I propose to work on three promising lines of investigation. 1- New architectures with joint training of WFSA and DNN parameters may bridge the performance gap when enough training data is available. 2 - New loss functions closer to actual sequence-based objective functions such as word or phoneme error rate should yield better performance than approximate losses used in deep neural only models. 3- Partial supervision afforded by structured generative models can significantly reduce the need for transcribed data. Although applicable to a wide range of problems, integrated models will have their largest impact where underlying structure is complex and annotated data is scarce. Thus I intend to apply them first to problems I encountered in my recent work on Indigenous languages spoken in Canada, ranging from subword analysis to speech recognition. Making speech technology accessible to these languages will benefit their transcription, preservation, and revitalization. This research addresses key limitations of current deep learning models in speech recognition, but potentially has broader applications in natural language processing, machine translation, or genomics, where sequence-to-sequence and segmentation problems are common. Because it combines the solid mathematical framework of probabilistic models with the practical, scalable methods of deep learning, this approach is well positioned to generate advances in knowledge while providing a rich learning environment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Deep structured speech models
-
批准号:DGECR-2021-00092
-
项目类别:Discovery Launch Supplement
-
资助金额:$0.91万
-
财政年份:2021
-
负责人:Boulianne, Gilles
-
依托单位:
Deep structured speech models
-
批准号:RGPIN-2021-02652
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.77万
-
财政年份:2021
-
负责人:Boulianne, Gilles
-
依托单位:
海外基金