课题基金 / 基金详情

Naturally Sounding Speech Synthesis and Recognition Based on the Formulation of Prosody

Naturally Sounding Speech Synthesis and Recognition Based on the Formulation of Prosody
基于韵律表述的自然语音合成与识别
批准号:
09480061
负责人:
HIROSE Keikichi
金额:
$5.25万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (B)
财政年份:
1997
资助国家:
日本
项目状态:
已结题
起止时间:
1997 至 1999

项目摘要

项目成果

HIROSE Keikichi的其他基金

相似基金

相关文献

中文摘要
翻译
本研究旨在阐明语音韵律特征与语言信息、准语言信息和非语言信息之间的关系,实现语音合成的先进技术:1.通过对基频轮廓进行低通滤波来抑制重音成分并提取其偏差值,从而提高了从基频轮廓中自动提取短语成分的精度。在自动韵律标注方法中实现了准确率的进一步提高,该方法利用从语言信息中获得的韵律知识作为约束F0参数估计。构造了类对话语音合成的Mora时长规则。这些规则基本上是将阅读风格的每个Mora时长修改为以韵律短语为基础的对话式演讲的时长,由FO轮廓定义。不同态度…语音的韵律特征更多的是对S/情绪的分析。研究发现,说话者有选择地控制几个韵律线索来表达态度/情感的程度。通过知觉实验还发现,分段特征控制对于实现情感话语也是不可或缺的。提出了一种以Mora为单位用代码表示韵律词的F0轮廓并对其转换进行统计建模的方法(Moraic转换的统计模型)。韵律词边界的插入误差为11%~15%,检测率为70%~75%。将该方法应用于连续语音识别,对Mora识别率的改善不大。提出了一种利用重音类型和短语边界位置的输入生成句子F0轮廓的方法。针对大词汇量连续语音识别波束搜索过程中的动态剪枝问题,提出了一种基于韵律特征的动态剪枝方法。实验证明,在不降低识别率的情况下,搜索空间可以减少到四分之一。该方法增大了韵律边界处的波束宽度,减小了边界之间的波束宽度。提出了一种利用韵律边界信息选择具有不同语境依赖关系的音素模型的方法。在此基础上,开发了一个学术信息检索口语对话系统,并对该系统进行了评价。较少
英文摘要
Several results including the following ones were achieved through the study aiming at formulating the relationship between prosodic features of speech and linguistic and para/non linguistic information, and realizing advanced technologies on speech synthesis :1. An improved accuracy was realized in automatic extraction of phrase component onsets from fundamental frequency (FO) contours by suppressing accent components through low-pass filtering of the contours and by taking their deviations. Further improvements in accuracy were realized in a method of automatic prosodic labeling where knowledge on prosody obtainable from linguistic information was utilized as constrictions F0 parameter estimation.2. Mora duration rules were constructed for dialogue-like speech synthesis. These rules are basically to modify each mora duration of reading-style speech to that of dialogue-like speech in prosodic phrase-basis, defined by the FO contours.3. Prosodic features of Speech with various attitude … More s/emotions were analyzed. It was found that a speaker selectively controlling several prosodic cues to express degree of attitudes/emotion. It was also found through a perceptual experiment that segmental feature control were also indispensable to realized emotional speech.4. A method was developed to represent F0 contours of prosodic words by codes in mora unit and to model their transitions statistically (Statistic model of moraic transition). The detection rates of 70-75% were achieved with insertion errors of 11-15% for prosodic word boundaries. The method was applied to continuous speech recognition with few % improvements in mora recognition rates. A method was also developed to generate sentence F0 contours with inputs of accent types and phrase boundary positions.5. A prosodic feature-based method was developed for the dynamic pruning in beam search process of large-vocabulary continuous speech recognition. It was proved that the search space could be reduced to a quarter without degradation in recognition rates. The method enlarges beam width at prosodic boundaries and decreases between boundaries. A method was also developed to select phoneme models with various context dependencies using prosodic boundary information.6. Based on the results obtained, a spoken dialogue system of academic information retrieval was developed and evaluated. Less
期刊论文(98)
专著(0)
科研奖励(0)
会议论文
桜井淳宏: "Detecting accent sandhi in Japanese using a superpositional F_0 model"Proc.European Conf,on Speech Communication and Technology. 4. 1863-1866 (1999)
Atsuhiro Sakurai:“使用叠加 F_0 模型检测日语连读重音”Proc.European Conf,关于语音通信和技术 4. 1863-1866 (1999)。
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
DOI: --
发表时间:
期刊:
影响因子: --
作者: []
通讯作者:
共 92 条
    Pronunciation education system based on the systematization of non-mothor tongue speech prosody using generation process model and speech synthesis
    • 批准号:
      24652115
    • 项目类别:
      Grant-in-Aid for Challenging Exploratory Research
    • 资助金额:
      $2.33万
    • 财政年份:
      2012
    • 负责人:
      HIROSE Keikichi
    • 依托单位:
    Advanced method of prosody control in statistical-based speech synthesis using generation process model of fundamental frequency contours
    • 批准号:
      24300068
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $11.4万
    • 财政年份:
      2012
    • 负责人:
      HIROSE Keikichi
    • 依托单位:
    Expressive Multi-language Speech Synthesis Based on the Generation Process Model and Its Use for Automatic Speech Translation
    • 批准号:
      21300061
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $11.23万
    • 财政年份:
      2009
    • 负责人:
      HIROSE Keikichi
    • 依托单位:
    Synthesis of speech in any speaking styles based on corpus-based generation of prosodic features using the generation process model
    • 批准号:
      17300055
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $10.79万
    • 财政年份:
      2005
    • 负责人:
      HIROSE Keikichi
    • 依托单位:
    海外基金