课题基金 / 基金详情

A Study on Constructing Various Acoustic Models using Distributed Speech Corpora

A Study on Constructing Various Acoustic Models using Distributed Speech Corpora
利用分布式语音语料库构建多种声学模型的研究
批准号:
15200014
负责人:
TAKEDA Kazuya
金额:
$29.37万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (A)
财政年份:
2003
资助国家:
日本
项目状态:
已结题
起止时间:
2003 至 2005

项目摘要

项目成果

TAKEDA Kazuya的其他基金

相似基金

相关文献

中文摘要
翻译
为了收集各种环境条件下的语音,对公共交通引导、车载信息检索和公共空间引导进行了语音对话系统的现场测试。基于这三个语料库,开发了声学模型训练数据共享基础架构的原型。在这个系统中,人们可以通过查询说话者的年龄、话语的信噪比和音素频率的分布来搜索特定的语音子集。该系统通过共享HMM声学模型各状态下的高斯混合模型的访问次数、分支次数、和平方和等有效统计量来训练一组HMM模型。此外,为了对话语进行特征化,即不需要显式语音活动检测(VAD),在较宽的信噪比范围内开发了信噪比方法。在训练策略方面,除了对话语集进行最大似然训练外,还研究了一种仅使用统计量的模型自适应方法。通过识别实验验证了基于预存储统计量的自适应方法的有效性,自适应训练的模型准确率几乎等同于混合EM算法。
英文摘要
In order to collect speech utterances made under various environmental conditions, field tests of spoken dialogue systems have been conducted for the public transportation guidance, the in-car information retrieval and the guidance for a public space. Based on the three corpora, a prototype of the data sharing infrastructure for acoustic model training has been developed. In the system, one can search for the particular speech subsets by invoking queries on the age of the speakers, SNR of the utterance and distribution of the phoneme frequency. The system can train a set of HMM's by sharing the efficient statistics, i.e., the visiting count, the branching count, the sum and the square sum, for the Gaussian Mixture pdf's for each state of HMM acoustic models. In addition, in order to characterize the utterance, a blind, i.e., does not require the explicit voice activity detection (VAD), method for SNR is developed for wide range of the SNR.As for the training strategy, not only the maximum likelihood (ML) training over the set of utterances, but also a model adaptation method using only statistics has been also studied. The effectiveness of the adaptation approach using pre-stored statistics for each utterance was confirmed through the recognition experiments where the accuracy of the model trained by the adaptation is almost equivalent to the pooled EM algorithm.
期刊论文(1163)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2005
期刊: Interspeech 2005
影响因子: --
作者: [Yasunori OHISHI, Masataka GOTO, Katunobu ITO, Kazuya TAKEDA]
通讯作者: Kazuya TAKEDA
局所的・大局的な特徴を利用した歌声と朗読音声の識別
使用局部和全局特征区分歌声和朗读声
DOI: --
发表时间: 2005
期刊: 情報処理学会 音楽情報科学研究会 Vol.2005・No.82
影响因子: --
作者: [大石康智, 後藤真孝, 伊藤克亘, 武田一哉]
通讯作者: 武田一哉
実走行車内単語音声データベースCENSREC-3と共通評価環境の構築
使用实际车载词汇数据库CENSREC-3构建通用评估环境
DOI: --
发表时间: 2005
期刊: 情報処理学会研究報告 2005-SLP-55(8)
影响因子: --
作者: [藤本雅清, 中村哲, 武田一哉, 黒岩眞吾, 山田武志, 北岡教英, 山本一公, 水町光徳, 西浦敬信, 佐宗晃, 宮島千代美, 遠藤俊樹]
通讯作者: 遠藤俊樹
ケプストラム分析を用いた室内伝達関数のモデル化の検討
考虑使用倒谱分析进行室内传递函数建模
DOI: --
发表时间: 2005
期刊: 電子情報通信学会技術研究報告 EA2004-139
影响因子: --
作者: [斉藤文訓, 西野隆典, 伊藤克亘, 武田一哉]
通讯作者: 武田一哉
451
    A study of comparative history and 3D archive creating on the walled cities and Buddist stupas in eastern Eurasia during the 5-13th century.
    • 批准号:
      18K00918
    • 项目类别:
      Grant-in-Aid for Scientific Research (C)
    • 资助金额:
      $2.83万
    • 财政年份:
      2018
    • 负责人:
      TAKEDA Kazuya
    • 依托单位:
    Interdisciplinary research on the use of historical and archaeological materials to build a diversity of crop resources for the next generation.
    • 批准号:
      18KT0048
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $11.81万
    • 财政年份:
      2018
    • 负责人:
      TAKEDA Kazuya
    • 依托单位:
    Interdisciplinary research on the distribution, cultivation, and food cultural history on the cruciferous crops
    • 批准号:
      26300003
    • 项目类别:
      Grant-in-Aid for Scientific Research (B)
    • 资助金额:
      $9.98万
    • 财政年份:
      2014
    • 负责人:
      TAKEDA Kazuya
    • 依托单位:
    Analysis of individuality in subjective similarity among music songs
    • 批准号:
      25540168
    • 项目类别:
      Grant-in-Aid for Challenging Exploratory Research
    • 资助金额:
      $2.41万
    • 财政年份:
      2013
    • 负责人:
      TAKEDA Kazuya
    • 依托单位:
    海外基金