课题基金 / 基金详情

Deep speech representation learning for research in phonetics

Deep speech representation learning for research in phonetics
用于语音学研究的深度语音表示学习
批准号:
446378607
负责人:
Professor Dr.-Ing. Reinhold Häb-Umbach
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:

项目摘要

项目成果

Professor Dr.-Ing. Reinhold Häb-Umbach的其他基金

相似基金

相关文献

中文摘要
翻译
语音信号是一种丰富的信息源,它不仅传达语言信息,而且传达语言外/副语言信息,例如说话人的身份,性别,情绪状态,年龄或社会地位。然而,这些特征隐藏在复杂的,不透明的语音信号的变化,大多数是模糊的语音研究。随着深度学习的出现,特别是深度生成建模技术的出现,语音合成和语音转换技术取得了长足的进步,我们认为合成语音可以成为语音学研究的重要工具。因此,本项目的总体目标是探索语音的深度生成建模作为支持语音学基础研究的工具的潜力。为了限制任务,我们将不考虑从文本中合成刺激,而是专注于语音的专用操作以生成具有所需属性的新语音信号。目标是开发生成模型,其通过潜变量提供语音信号的表示,其是紧凑的并且提供关于观察到的语音信号的信息,其通过表示的不同维度表示语音信号的不同变化源,其允许沿着沿着语音上合理的维度对语音提示进行专用操纵,并且其服从于人类解释。有了这些工具,语音学家可以控制低层次的声学语音属性和高层次的抽象概念。这里考虑的高层次概念将是说话者和语言内容相关的语音信号的变化,以及方言线索的隔离的解开和专门的操纵。作为数据驱动的,开发的技术将,但是,有可能是有用的,也为其他额外的或非语言特征的研究,提供适当的训练数据。所开发的工具的性能和效用将通过机器和人类感知实验以及语音专家对语音可识别性的审查来衡量。
英文摘要
The speech signal is a rich source of information that conveys not only linguistic but also extra/para-linguistic information, such as the speaker's identity, gender, emotional state, age, or the social status. However, those traits are hidden in complex, non-transparent variations of the speech signal, and mostly obscure to speech research. With recent progress in speech synthesis and voice conversion caused by the advent of deep learning, notably by deep generative modeling, we argue that synthesized speech can become a valuable tool for research in phonetics.The overarching goal of this project is thus to explore the potential of deep generative modeling of speech as a tool to support basic research in phonetics. To constrain the task, we will not consider the synthesis of stimuli from text, but concentrate on the dedicated manipulation of speech to generate new speech signals with desired properties. The goal is to develop generative models which offer a representation of the speech signal by latent variables, which is compact and informative about the observed speech signal, which represents different sources of variation of the speech signal by different dimensions of the representation, which allows a dedicated manipulation of a phonetic cue along phonetically plausible dimensions, and which is amenable to human interpretation. With these tools a phonetician will be given control over both low-level acoustic-phonetic properties as well as high-level abstract concepts. The high-level concepts considered here will be the disentanglement and dedicated manipulation of speaker and linguistic content related variations of a speech signal, as well as the isolation of dialectal cues. Being data-driven, the developed techniques will, however, have the potential to be useful also for the study of other extra- or paralinguistic traits, provided appropriate training data is available. The performance and utility of the developed tools will be measured by both machine and human perception experiments, as well as by phonetic expert scrutiny for phonetic plausibility.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Coordination Funds
Source separation and noise reduction for automatic speech recognition in dynamic acoustic scenarios
Sound recognition with limited supervision over sensor networks
  • 批准号:
    318489874
  • 项目类别:
    Research Units
  • 资助金额:
    $0.0万
  • 财政年份:
    2016
  • 负责人:
    Professor Dr.-Ing. Reinhold Häb-Umbach
  • 依托单位:
Bayesian Learning of a Hierarchical Representation of Language from Raw Speech
国内基金
海外基金
儿童植入耳蜗后听觉行为与言语发展进程的关联性研究
  • 批准号:
    81170916
  • 项目类别:
    面上项目
  • 资助金额:
    65.0万元
  • 批准年份:
    2011
  • 负责人:
    刘莎
  • 依托单位:
儿童植入人工耳蜗后开放式听觉言语发育特性研究
  • 批准号:
    30872859
  • 项目类别:
    面上项目
  • 资助金额:
    30.0万元
  • 批准年份:
    2008
  • 负责人:
    刘莎
  • 依托单位: