A Study for Utilizing the Linguistic Information in Phoneme Recognition to Understand Continuous Speech
A Study for Utilizing the Linguistic Information in Phoneme Recognition to Understand Continuous Speech
批准号:
03452173
负责人:
KIDO Ken'iti
金额:
$4.35万
依托单位国家:
日本
项目类别:
Grant-in-Aid for General Scientific Research (B)
财政年份:
1991
资助国家:
日本
项目状态:
已结题
起止时间:
1991 至 1993
中文摘要
在本研究中,我们提出了两种较高性能的音素识别方法和利用目标音素周围语言信息的连续语音识别方法。首先,我们提出了基于小波变换的MR-HMM(多分辨率HMM),它能够控制时频分辨率。提出了用小波变换树数据来表示小波变换得到的尺度图的时频空间。利用这种WTD结构,我们提出了基于MR-HMM的状态合并算法,实现了较高的识别率。其次,除了最常用但还不够充分的倒谱参数外,我们还提出了利用9个声学特征进行音素识别的方法。一般来说,需要使用几种声学参数来分析哪些参数适合于指定的音素识别。但是,所提出的方法允许使用除以下几种参数之外的多种参数。我们提出了隶属度的概念,将线性判别方法,即两类判别应用于多类判别。最后,本文提出了一种新的语言识别方法,该方法利用句子中词语的共现关系进行识别。这种方法不使用语法知识,因此语音识别任务是可行的。将该语言识别方法与上述声学识别方法相结合,可以通过语言识别阶段来控制声学识别阶段的误识别。实验结果表明,本文提出的识别方法是有效的。
英文摘要
In this study, we proposed 2 higher performance phoneme recognition methodsand the continuous speech recognition method utilizing the linguistic information around the target phoneme.At first, we proposed MR-HMM (Multi-Resolution HMM) based on Wavelet transform, which is able to control the time-frequency resolution. The WTD (Wavelet transform Tree Data) is proposed to represent the time-frequency space in scalogram that is obtained through Wavelet transform. Using this WTD structure, we proposed the State merge Algorithm stucying MR-HMM, it enables the high recognition rate.Next, we proposed the phoneme recognition method using the 9 acoustic features besides the cepstrum parameters that is most popular but not enough. In general, it is necessary for using the several kinds of acoustic parameters to analyze what parameters are suitable for the specified phoneme recognition. But, the proposed method enables using the several kinds of parameters except that. We proposed the Membership Scale to enable applying the linear discriminant method that is for 2 category discrimination to the multi category discrimination. Using this method, the linguistic recognition stage can get the reliability of the results from the acoustical recognition stage.Finally, we proposed the new linguistic recognition method, that uses the co-occurative relationship of the words in one sentence. This method doesn't use the grammatical knowledge, so the task fre speech is available. Combining this linguistic recognition method with the acoustic recognition methods mentioned above, the misrecognition in the acoustical recognition stage can be controlled by the linguistic rrecognition stage. From the experimental results, we confirmed the effectiveness of the proposed recognition methods.
期刊论文(20)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
小林淳: "動詞、名詞のスポッティングによる会話文の認識" 日本音響学会秋季研究発表会講演論文集. 175-176 (1993)
Jun Kobayashi:“通过识别动词和名词来识别会话句子”日本声学学会秋季研究会议记录 175-176(1993)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
荒井 秀一: "A Network for Phenome Recognition by Spectral Local Peaks" Proc.14th International Congress on Acoustics. G-4-1. 877-878 (1992)
Shuichi Arai:“通过光谱局部峰进行现象组识别的网络”Proc.14th 国际声学大会 G-4-1(1992 年)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
棚橋健二: "正常および異常音声のフォルマント周波数の時間遷移パターンによる比較" 日本音響学会秋期研究発表会講演論文集. 595-596 (1993)
Kenji Tanahashi:“基于时间转换模式的正常和异常语音的共振峰频率的比较”日本声学学会秋季研究会议论文集 595-596(1993)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
栗原世治: "各種音響パラメータが保持する個人性情報の分析" 日本音響学会秋季研究発表会講演論文集. 645-646 (1993)
Seiji Kurihara:“各种声学参数所持有的个人信息的分析”日本声学学会秋季研究会议记录 645-646(1993)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
大内康裕: "正常および異常音声の第1・第2フォルマント平面における比較" 日本音響学会秋期研究発表会講演論文集. 593-594 (1993)
Yasuhiro Ouchi:“第一和第二共振峰平面中正常和异常语音的比较”日本声学学会秋季研究会议论文集 593-594(1993)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 13 条
RESEARCH ON THE DEVELOPMENT OF SPEECH QUALITY EVALUATION SYSTEM FOR THE PURPOSE OF EDUCATION AND TRAIN
-
批准号:05555104
-
项目类别:Grant-in-Aid for Developmental Scientific Research (B)
-
资助金额:$10.24万
-
财政年份:1993
-
负责人:KIDO Ken'iti
-
依托单位:
Developmental Research of an automatic inspector machine for diagnosis of each part of roll bearings based on vibration analysis
-
批准号:63850049
-
项目类别:Grant-in-Aid for Developmental Scientific Research
-
资助金额:$7.55万
-
财政年份:1988
-
负责人:KIDO Ken'iti
-
依托单位:
A study on the conkersion from a sentence speech to a kanji-kana string using phoneme recognition, syntax and semantics processings
-
批准号:59420031
-
项目类别:Grant-in-Aid for General Scientific Research (A)
-
资助金额:$14.14万
-
财政年份:1984
-
负责人:KIDO Ken'iti
-
依托单位:
海外基金