Large Vocabulary Continuous Speech Recognition System on Japanese Newspaper Reading Task
Large Vocabulary Continuous Speech Recognition System on Japanese Newspaper Reading Task
批准号:
10680368
负责人:
KOHDA Masaki
金额:
$2.11万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1998
资助国家:
日本
项目状态:
已结题
起止时间:
1998 至 2000
中文摘要
本文对大词汇量连续语音识别(LVCSR)系统在日语报纸阅读任务中的应用进行了研究,得到了以下结果。(1)声学模型:隐马尔可夫网络(hmm - net)是一种高度精确和鲁棒的声学模型,它将上下文相关隐马尔可夫模型作为一个网络表示为一个固定状态结构。提出了一种基于状态聚类的快速拓扑设计方法,用于LVCSR生成高精度的hmm - nets。在此基础上,研究了基于最大似然线性回归(MLLR)的说话人自适应声学模型,并提出了一种基于BIC原理的回归类选择算法。(2)语言模型:研究了N-gram任务自适应方法,该方法使用大的一般任务语料库(TI文本)和小的特定任务语料库(AD文本),采用简单的加权混合TI和AD文本。此外,我们提出了一个新的SCFG(随机上下文自由语法)模型,该模型使用基于短语的依赖语法…More r代替一般的CFG。使用基于SCFG模型和三元图的混合模型的错误率比只使用三元图的错误率小。(3)解码器:研究了LVCSR的快速搜索策略,提出了一种基于音素图的假设约束方法,有效地对搜索空间进行了修剪。该方法在预处理阶段生成一个音素图,然后在主识别阶段利用音素图信息限制假设扩展的同时搜索最佳单词序列。在使用词图作为中间数据结构的多通道LVCSR系统中,为了生成好的词图,需要对解码器参数进行优化。提出了一种新的参数优化方法。该方法使用双字LM对词图进行重新评分,而不是为每个参数设置生成许多词图。(4)软件工具:我们描述了一个基于词和类的n-gram的统计语言模型工具包。该工具包与CMU-Cambridge SLM工具包具有命令级兼容性,并支持类n-gram和n-gram计数混合以及使用线性插值的组合语言模型。少
英文摘要
We investigated large vocabulary continuous speech recognition (LVCSR) system on Japanese newspaper reading task, and obtained the following results.(1) Acoustic models : A Hidden Markov Network (HM-Net) is a highly accurate and robust acoustic model which represents a tied-state structure of context dependent Hidden Markov Models as a network. We propose a state clustering-based rapid topology design method to generate high accuracy HM-Nets for LVCSR.Furthermore, MLLR (Maximum Likelihood Linear Regression)-based speaker adaptation of acoustic models is investigated, and a regression class selection algorithm based on the BIC principle is proposed.(2) Language models : N-gram task adaptation method is investigated, which uses large corpus of the general task (TI text) and small corpus of the specific task (AD text), and employs a simple weighting to mix TI and AD texts. Furthermore we propose a new SCFG (Stochastic Context Free Grammar) model which uses a phrase-based dependency gramma … More r instead of general CFG.Word error rate in the case of using the mixture model besed on the proposed SCFG model and trigram becomes less than that in the case of using only the trigram.(3) Decoder : We investigate about fast search strategies for LVCSR, and propose a new method - a phoneme-graph-based hypothesis restriction, which effectually prunes the search space. In the proposed method, a phoneme graph is generated at the pre-processing stage, and then the best word sequence is searched while restricting expansion of hypotheses using the information of the phoneme graph at the main recognition stage. In the multiple pass LVCSR system that uses word graph as an intermediate data structure, decoder parameters should be optimized in order to generate a good word graph. A new method to optimize these parameters is proposed. This method uses rescoring of the word graph using bigram LM instead of generating many word graphs for each parameter setting.(4) Software Tool : We describe a statistical language model toolkit for word and class-based n-gram. This toolkit has command-level compatibility with CMU-Cambridge SLM Toolkit, and supports class n-gram and n-gram count mixture as well as combined language model using linear interpolation. Less
期刊论文(49)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
A.Ito, M.Kohda, M.Ostendorf: "A New Metric for Stochastic Language Model Evaluation"Proc. Euro. Conf. on Speech Commu. and Technology. Vol.4. 1591-1594 (1999)
A.Ito、M.Kohda、M.Ostendorf:“随机语言模型评估的新指标”Proc。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
加藤正治: "単語グラフ生成におけるパラメータ最適化の検討"電子情報通信学会技術研究報告. SP2000-93. 107-112 (2000)
加藤正治:“字图生成中的参数优化研究”IEICE技术研究报告107-112(2000)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
伊藤彰則: "単語およびクラスN-gram作成のためのツールキット"電子情報通信学会技術研究報告. SP2000-106. 67-72 (2000)
Akinori Ito:“创建单词和类别 N 元语法的工具包”IEICE SP2000-106 (2000)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
堀 智織: "確率文脈自由文法を用いた言語モデルの構築と音声認識実験による評価"電子情報通信学会技術研究報告. SP99-37. 79-86 (1999)
Tomoori Hori:“使用概率上下文无关语法构建语言模型并通过语音识别实验进行评估”IEICE 技术报告 SP99-37 (1999)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
斎藤 秀樹: "bigram に基づく ergodic HMM による言語モデルの検討"日本音響学会講演論文集. 3-1-3. 101-102 (1999)
Hideki Saito:“基于二元语法的遍历 HMM 的语言模型研究”日本声学学会论文集 3-1-3(1999)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 49 条
Large-vocabulary continuous speech recognition on spontaneous speech task
-
批准号:18500126
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.22万
-
财政年份:2006
-
负责人:KOHDA Masaki
-
依托单位:
Spontaneous speech recognition
-
批准号:15500098
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.05万
-
财政年份:2003
-
负责人:KOHDA Masaki
-
依托单位:
Algorithm of Spontaneous Speech Recognition Based on A^<**> Search
-
批准号:07680379
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.09万
-
财政年份:1995
-
负责人:KOHDA Masaki
-
依托单位:
Speech Recognition Based on Intelligent Beam Search Algorithm
-
批准号:01460254
-
项目类别:Grant-in-Aid for General Scientific Research (B)
-
资助金额:$4.42万
-
财政年份:1989
-
负责人:KOHDA Masaki
-
依托单位:
海外基金