Continuous speech recognition with adaptabilty to the speaking rate of an input speech
Continuous speech recognition with adaptabilty to the speaking rate of an input speech
批准号:
07458064
负责人:
MAKINO Shozo
金额:
$4.1万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (B)
财政年份:
1995
资助国家:
日本
项目状态:
已结题
起止时间:
1995 至 1997
中文摘要
本研究开发了一个语音识别系统,该系统使用从输入语音的语速估计的音素持续时间信息。在本研究中,说话速率被假设为反映到平均元音长度。声学处理器使用修改的LVQ 2将输入语音变换成相似性矩阵。根据初步识别结果计算平均元音长度。根据输入语音中元音的平均长度估计每个单词模板中每个音素的持续时间。在考虑音素时长估计的基础上,使用DTW进行了口语词识别实验,对5名男性说话人的212个词汇(测试集)的词识别率为97.3%。音素持续时间信息是从另外5名男性和10名女性说话者发出的212个单词词汇(训练集)中收集的。其中,先用音素相关估计和后用音素相关估计的混合组合效果最好,并将上述方法推广到音素识别中。对5名男性使用者的212个词汇的音素识别率从71.8%提高到86.3%。
英文摘要
This tesearch developed a spoken word recognition system which used phoneme duration information estimated from the speaking rate of an input speech. In this research, the speaking rate is assumed to be reflected to the average vowel length. The acoustic processor transforms the input speech into a similarity matrix using the modified LVQ2. The average vowel length is computed from the preliminary recognition result. The duration of each phoneme in each word template is estimated from the average length of vowels in the input speech. By taking into account the estimated phoneme duration, the spoken word recognition experiments were carried out using the DTW.The word recognition score was 97.3% for the 212 word vocabulary uttered by 5 male speakers (test set). The phoneme duration information is collected from the 212 word vocabulary uttered by another 5 male and 10 female speakers (training set). The hybrid combination of the prceiding phoneme dependent estimation and the follwoing phoneme dependent estimation gave the best performance.The above-mentioned method was extended to phoneme recognition. The phoneme accuracy increased from 71.8% to 86.3% for phonemes in the 212 word vocabulary uttered by 5 male speakers (test set).
期刊论文(22)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
原田, 鈴木, 牧野: "離散型HMnetによる新聞記事からの文節モデルの獲得" 電子情報通信学会技術報告. SP97・24. 45-50 (1997)
Harada、Suzuki、Makino:“使用离散 HMnet 从报纸文章中获取短语模型”SP97·24 (1997)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
沖本,牧野: "可変長パターンと識別学習を用いた音素認識" 信学技報. Vol. 96 No. 93. 7-14 (1996)
Okimoto, Makino:“使用可变长度模式和判别性学习的音素识别”,IEICE 技术报告,第 96 卷,第 93 期。7-14 (1996)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
M.SUZUKI,S.MAKINO,A.ITO,H.ASO,H.SHIMODAIRA: "A New HMnet Construction Algorithm Requiring No Contextual Factors" IEICE Trans.on Information and Systems. E78-D,6. 662-668 (1995)
M.SUZUKI、S.MAKINO、A.ITO、H.ASO、H.SHIMODAIRA:“一种不需要上下文因素的新 HMnet 构建算法”IEICE Trans.on 信息和系统。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
S.MAKIKO, M.SUZUKI, A.HARADA: "Automatic Acquistion of Language Model using HMnet" Proc.Int.Conf Speech Processing'97. I. 47-54 (1997)
S.MAKIKO、M.SUZUKI、A.HARADA:“使用 HMnet 自动获取语言模型”Proc.Int.Conf 语音处理97。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
大坂,牧野: "発声速度に基づく音素持続時間予測を用いた音素認識" 信学技報. Vol. 96 No. 93. 1-6 (1996)
Osaka, Makino:“基于语速的音素持续时间预测的音素识别”IEICE 技术报告,第 96 卷第 93. 1-6 (1996)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 22 条
Japanese text dictation system for official reports
-
批准号:07558042
-
项目类别:Grant-in-Aid for Scientific Research (B)
-
资助金额:$3.39万
-
财政年份:1995
-
负责人:MAKINO Shozo
-
依托单位:
Study on utilization and effectiveness of linguistic information in the word recognition based on phoneme, syllable or character sequence with errors
-
批准号:63460222
-
项目类别:Grant-in-Aid for General Scientific Research (B)
-
资助金额:$4.16万
-
财政年份:1988
-
负责人:MAKINO Shozo
-
依托单位:
Shape Estimation and Detection of Defects of a Structural Object from Acoustic Signal Using Digital Signal Processing and Intellectual Processing
-
批准号:63420037
-
项目类别:Grant-in-Aid for General Scientific Research (A)
-
资助金额:$7.81万
-
财政年份:1988
-
负责人:MAKINO Shozo
-
依托单位:
海外基金