课题基金 / 基金详情

Deep structured speech models

Deep structured speech models
深层结构化语音模型
批准号:
RGPIN-2021-02652
负责人:
Boulianne, Gilles
金额:
$2.77万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31

项目摘要

项目成果

Boulianne, Gilles的其他基金

相似基金

相关文献

中文摘要
翻译
最先进的语音识别模型是端到端模型或混合模型。端到端模型完全基于深度神经网络(DNN)。混合模型联合收割机将加权有限状态自动机(WFSA)的底层结构与深度神经网络的表层相结合。当在足够多的类似于测试数据的注释记录上进行训练时,端到端模型可以优于混合模型。对于不匹配或较小的训练数据,混合模型通常是更好的选择,因为它们可以使用编码在其结构中的先验语言知识来更好地泛化并避免过拟合。然而,在混合模型中,有限状态自动机和深度神经组件没有集成;它们是用不同的目标和算法独立训练的。因此,即使有足够的训练数据,性能差距仍然存在。在实践中,混合模型训练起来很复杂,需要专门的编码,并且无法轻松地与Pytorch或Tensorflow等常见的深度学习框架集成。因此,它们的实现落后于一般深度学习的最新发展。最近提出的可微分自动机建议如何将它们与深度神经网络集成到一个具有端到端可微分损失的单一模型中,以一种在时间和空间上高效的方式来解决语音问题。这为语音建模开辟了几个几乎尚未探索的研究领域。我建议从三个有希望的调查方向着手。1-具有WFSA和DNN参数的联合训练的新架构可以在足够的训练数据可用时弥合性能差距。2.新的损失函数更接近于实际的基于序列的目标函数,如单词或音素错误率,应该比仅深度神经模型中使用的近似损失产生更好的性能。3-结构化生成模型提供的部分监督可以显着减少对转录数据的需求。尽管综合模型适用于各种问题,但在基础结构复杂和注释数据稀缺的情况下,综合模型将产生最大的影响。因此,我打算首先将它们应用于我最近在加拿大土著语言工作中遇到的问题,从子词分析到语音识别。使这些语言能够使用语音技术将有利于它们的转录,保存和振兴。这项研究解决了当前深度学习模型在语音识别中的主要局限性,但可能在自然语言处理、机器翻译或基因组学中有更广泛的应用,其中序列到序列和分割问题很常见。由于它将概率模型的坚实数学框架与深度学习的实用、可扩展方法相结合,因此这种方法能够在提供丰富学习环境的同时促进知识进步。
英文摘要
State-of-the-art speech recognition models are either end-to-end models or hybrid models. End-to-end models are entirely based on deep neural networks (DNN). Hybrid models combine an underlying structure of weighted finite-state automata (WFSA) with a surface layer of deep neural networks. When trained on large enough amounts of annotated recordings similar to the test data, end-to-end models can outperform hybrid models. For mismatched or smaller training data, hybrid models are often a better choice since they can use prior linguistic knowledge encoded in their structure to better generalize and avoid overfitting. However, in hybrid models, finite-state automata and deep neural components are not integrated; they are trained independently with different objectives and algorithms. As a result, a performance gap remains even when enough training data is available. In practice, hybrid models are complicated to train, require specialized coding, and cannot be easily integrated with common deep learning frameworks such as Pytorch or Tensorflow. Thus their implementations lag behind the latest developments in general deep learning. Recent proposals for differentiable automata suggest how they could be integrated with deep neural networks into a single model with an end-to-end differentiable loss, in a way that is efficient in time and space to be scalable enough for speech problems. This opens up several areas of research which are yet almost unexplored for speech modelling. I propose to work on three promising lines of investigation. 1- New architectures with joint training of WFSA and DNN parameters may bridge the performance gap when enough training data is available. 2 - New loss functions closer to actual sequence-based objective functions such as word or phoneme error rate should yield better performance than approximate losses used in deep neural only models. 3- Partial supervision afforded by structured generative models can significantly reduce the need for transcribed data. Although applicable to a wide range of problems, integrated models will have their largest impact where underlying structure is complex and annotated data is scarce. Thus I intend to apply them first to problems I encountered in my recent work on Indigenous languages spoken in Canada, ranging from subword analysis to speech recognition. Making speech technology accessible to these languages will benefit their transcription, preservation, and revitalization. This research addresses key limitations of current deep learning models in speech recognition, but potentially has broader applications in natural language processing, machine translation, or genomics, where sequence-to-sequence and segmentation problems are common. Because it combines the solid mathematical framework of probabilistic models with the practical, scalable methods of deep learning, this approach is well positioned to generate advances in knowledge while providing a rich learning environment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Deep structured speech models
Deep structured speech models
海外基金