课题基金 / 基金详情

Generative Kernels and Score Spaces for Classification of Speech

Generative Kernels and Score Spaces for Classification of Speech
用于语音分类的生成核和评分空间
批准号:
EP/I006583/1
负责人:
Mark Gales
金额:
$49.96万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2011
资助国家:
英国
项目状态:
已结题
起止时间:
2011 至 --

项目摘要

项目成果

Mark Gales的其他基金

相似基金

相关文献

中文摘要
翻译
该项目的目标是显著提高自动语音识别系统在各种环境、说话者和说话风格下的性能。最先进的语音识别系统的性能在相当可控的条件下和背景噪声水平较低的情况下通常是可以接受的。然而,在许多现实情况下,可能存在高水平的背景噪音,例如车内导航,或广泛的频道条件和说话风格,例如在youtube上观察到的数据。语音识别系统的脆弱性是语音识别系统没有得到更广泛部署和使用的主要原因之一。它限制了语音可以可靠使用的可能领域,并且增加了开发应用程序的成本,因为必须调整系统以限制这种脆弱性的影响。这包括收集特定于领域的数据和对应用程序本身进行重大调优。绝大多数的语音识别研究都集中在提高基于隐马尔可夫模型(HMM)的系统的性能上。hmm是生成模型的一个例子,目前在最先进的语音识别系统中使用。在扬声器和噪声变化下,已经开发了许多方法来改善这些系统的性能。尽管有这些方法,系统还不够健壮,无法让语音识别系统达到界面自然度应该允许的影响水平。该项目将把语音社区中开发的当前生成模型与语音和机器学习社区中使用的判别分类器结合起来。提出的方法的一个重要的、新颖的方面是,生成模型被用来定义一个分数空间,该分数空间可以被判别分类器用作特征。这种方法有很多优点。可以使用当前最先进的自适应和鲁棒性方法来补偿特定扬声器和噪声条件的声学模型。除了能够将这些方法的任何进展纳入方案之外,没有必要开发使判别分类器适应说话人、风格和噪声的方法。语音识别的主要问题之一是必须对可变长度的数据序列进行分类。使用生成模型还允许在不改变判别分类器的情况下处理语音数据的动态方面。最后一个优势是生成模型获得的分数空间的性质。像hmm这样的生成模型具有潜在的条件独立性假设,虽然使它们能够有效地表示数据序列,但不能准确地表示语音等数据序列中的依赖关系。与生成模型相关的得分空间不具有与原始生成模型相同的条件独立性假设。这允许对语音数据中的依赖关系进行更精确的建模。生成和判别分类器的组合将在当前系统表现不佳的两种非常困难的数据形式上进行研究。第一个任务是语音的不利环境识别。在这些情况下,有非常高水平的背景噪声,导致系统性能严重下降。本任务感兴趣的数据将与东芝欧洲研究有限公司合作指定。第二个感兴趣的任务是从广泛的说话风格和条件的大词汇量语音识别数据。谷歌提供了来自YouTube的转录数据,以便在高度多样化的数据上对系统进行评估。与目前最先进的方法相比,该项目将为这两项任务带来显著的性能提升。
英文摘要
The aim of this project is to significantly improve the performance of automatic speech recognition systems across a wide-range of environments, speakers and speaking styles. The performance of state-of-the-art speech recognition systems is often acceptable under fairly controlled conditions and where the levels of background noise are low. However for many realistic situations there can be high levels of background noise, for example in-car navigation, or widely ranging channel conditions and speaking styles, such as observed on YouTube-style data. This fragility of speech recognition systems is one of the primary reasons that speech recognition systems are not more widely deployed and used. It limits the possible domains in which speech can be reliably used, and increases the cost of developing applications as systems must be tuned to limit the impact of this fragility. This includes collecting domain specific data and significant tuning of the application itself.The vast majority of research for speech recognition has concentrated on improving the performance of hidden Markov model (HMM) based systems. HMMs are an example of a generative model and are currently used in state-of-the-art speech recognition systems. A wide number of approaches have been developed to improve the performance of these systems under speaker and noise changes. Despite these approaches, systems are not sufficiently robust to allow speech recognition systems to achieve the level of impact that the naturalness of the interface should allow. This project will combine the current generative models developed in the speech community with discriminative classifiers used in both the speech and machine learning communities. An important, novel, aspect of the proposed approach is that the generative models are used to define a score-space that can be used as features by the discriminative classifiers. This approach has a number of advantages. It is possible to use current state-of-the-art adaptation and robustness approaches to compensate the acoustic models for particular speakers and noise conditions. As well as enabling any advances in these approaches to be incorporated into the scheme, it is not necessary to develop approaches that adapt the discriminative classifiers to speakers, style and noise. One of the major problems with speech recognition is that variable length data sequences must be classified. Using generative models also allows the dynamic aspects of speech data to be handled without having to alter the discriminative classifier. The final advantage is the nature of the score-space obtained from the generative model. Generative models such as HMMs have underlying conditional independence assumptions that, whilst enabling them to efficiently represent data sequences, do not accurately represent the dependencies in data sequences such as speech. The score-space associated with a generative model does not have the same conditional independence assumptions as the original generative model. This allows more accurate modelling of the dependencies in the speech data.The combination of generative and discriminative classifiers will be investigated on two very difficult forms of data that current systems perform badly on. The first task is adverse environment recognition of speech. In these situations there are very high levels of background noise which causes severe degradation in system performance. Data of interest for this task will be specified in collaboration with Toshiba Research Europe Ltd. The second task of interest is large vocabulary speech recognition of data from a wide-range of speaking styles and conditions. Google has supplied transcribed data from YouTube to allow evaluation of systems on highly diverse data. The project will yield significant performance gains over current state-of-the-art approaches for both tasks.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
System Combination with Log-Linear Models
系统与对数线性模型的组合
DOI: --
发表时间: 2016
期刊:
影响因子: --
作者: [Yang J]
通讯作者: Yang J
DOI: 10.1109/icassp.2014.6854215
发表时间: 2014-05
期刊: 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者: [Jingzhou Yang;R. V. Dalen;Shi-Xiong Zhang;M. Gales]
通讯作者: Jingzhou Yang;R. V. Dalen;Shi-Xiong Zhang;M. Gales
DOI: --
发表时间:
期刊:
影响因子: --
作者: [Mark Gales (Author)]
通讯作者: Mark Gales (Author)
Annotating large lattices with the exact word error
用精确的单词错误注释大格子
DOI: --
发表时间: 2015
期刊:
影响因子: --
作者: [Van Dalen R C]
通讯作者: Van Dalen R C
共 6 条
    Multimodal Video Search by Examples
    • 批准号:
      EP/V006223/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $88.52万
    • 财政年份:
      2021
    • 负责人:
      Mark Gales
    • 依托单位:
    海外基金