A Study of Distributed Speaker Recognition Methods
A Study of Distributed Speaker Recognition Methods
批准号:
14350204
负责人:
KUROIWA Shingo
金额:
$5.89万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (B)
财政年份:
2002
资助国家:
日本
项目状态:
已结题
起止时间:
2002 至 2004
中文摘要
本文主要研究分布式说话人识别方法(DSR)。DSR将识别的结构和计算组件分为两个组件——终端上的前端处理和服务器上的说话人识别引擎。DSR最重要的优点是它可以使用高频成分,其中显示说话者特定的信息。另一方面,DSR必须压缩发送数据,以建立一个较低的比特率进行传输。为了实现高精度和低比特率,我们开发了以下技术。1)一种实时去偏方法,提高了对增加量化失真和识别误差的卷积噪声的鲁棒性。2)基于直方图的说话人模型和大地移动者距离的非参数说话人识别方法。这些方法在低比特率下建立了相当于麦克风输入的高精度,4.8kbps是欧洲电信标准协会(ETSI)标准分布式语音识别系统推荐的比特率。我们还提出了以下技术,不仅可以在DSR中使用,也可以在传统的电话网络中使用。3)基于音素的说话人识别方法,该方法在客户端对语音信号进行分割,选择最有效的说话人识别音素发送给服务器。4)基于语音合成的语音编解码快速声学模型自适应技术。5)基于mft的语音识别和基于hmm的语音合成的丢包隐藏算法。此外,我们在两年内每周记录四个人的声音,以调查声音特征的变化。现在,我们正在利用这些数据探索说话人的基本语音属性。
英文摘要
In this research, we focused on a Distributed Speaker Recognition Method (DSR). DSR separates the structural and computational components of recognition into two components - the front-end processing on the terminal and the speaker recognition engine on the server. The most important advantage of DSR is that it can use a high frequency component in which speaker-specific information is revealed. On the other hand, DSR has to compress the sending data to establish a lower bit rate for transmission. In order to achieve both high accuracy and low bit rate, we have developed the following techniques.1)A Real-time bias removal method that improves the robustness against convolutional noises, which increases quantization distortion and recognition error.2)A Nonparametric speaker recognition method that consists of a histogram-based speaker model and Earth Mover's Distance.These proposed methods have established a high accuracy equivalent to microphone input under the condition of a low bit rate, 4.8kbps, which is the bit rate recommended as the European Telecommunication Standards Institute (ETSI) Standard Distributed Speech Recognition System.We also proposed the following techniques that can be used not only in DSR but in conventional telephone networks.3)A Phoneme-dependent speaker recognition method in which speech signals are segmented at the client and the most effective phonemes for speaker recognition are selected to be sent to the server.4)A Rapid acoustic model adaptation technique for codec speech using speech synthesis.5)Packet-loss concealment algorithms using MFT-based speech recognition and HMM-based speech synthesis.Furthermore, we have been recording four peoples' voices every week over two years to investigate change in voice characteristics. Now, we are exploring essential voice attributes that characterize a speaker using these data.
期刊论文(193)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Earth Mover's Distanceを用いた分散型話者認識手法
利用推土机距离的分布式说话人识别方法
DOI:
--
发表时间:
2004
期刊:
日本音響学会2004年秋季研究発表会講演論文集 1
影响因子:
--
作者:
[D.V.Rao, Y.Sasaki, T.Yuasa, T.Akatsuka, 柘植覚]
通讯作者:
柘植覚
Robust Feature Extraction in a Variety of Input Devices on the Basis of ETSI Standard DSR Front-end
基于ETSI标准DSR前端的多种输入设备的鲁棒特征提取
DOI:
--
发表时间:
2002
期刊:
7th International Conference on Spoken Language Processing (ICSLP2002)
影响因子:
--
作者:
[K.Z.Liu, M.Q.Dao, T.Inoue, Satoru Tsuge]
通讯作者:
Satoru Tsuge
Koji Tanaka: "An acoustic model adaptation using HMM-based speech synthesis"Proceedings of IEEE International Conference on Natural Language Processing and Knowledge Engineering. Vol.1. 368-376 (2003)
Koji Tanaka:“使用基于 HMM 的语音合成的声学模型自适应”IEEE 国际自然语言处理和知识工程会议论文集。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Evaluation of frequency characteristic normalization method with multiple reference cepstrum on the Japanese newspaper article sentences speech corpus,
多参考倒谱频率特性归一化方法对日本报纸文章句子语音语料库的评价,
DOI:
--
发表时间:
2004
期刊:
Proceedings of the third International Conference on Information, November 29-December 2,2004,Tokyo, Japan
影响因子:
--
作者:
[Satoru Tsuge, Shingo Kuroiwa, Masami Shishibori, Kenji Kita, Fuji Ren]
通讯作者:
Fuji Ren
統計的手法を用いた音声信号の復元手法の改良
基于统计方法的音频信号恢复方法的改进
DOI:
--
发表时间:
2005
期刊:
徳島大学工学部研究報告 No.50
影响因子:
--
作者:
[黒岩眞吾]
通讯作者:
黒岩眞吾
共 81 条
On the physical factors which makes the mother tongue dialogues smoothly - through the comparison with the non-mother tongue
-
批准号:24650075
-
项目类别:Grant-in-Aid for Challenging Exploratory Research
-
资助金额:$2.5万
-
财政年份:2012
-
负责人:KUROIWA Shingo
-
依托单位:
Robust Speaker Recognition with Intra-Speaker Variability Compensation based on Long-Term Recorded Speech Corpus
-
批准号:21300060
-
项目类别:Grant-in-Aid for Scientific Research (B)
-
资助金额:$11.48万
-
财政年份:2009
-
负责人:KUROIWA Shingo
-
依托单位:
Analysis of Intra-Speaker Variation and Development of Distributed Speaker Recognition System
-
批准号:17300065
-
项目类别:Grant-in-Aid for Scientific Research (B)
-
资助金额:$10.65万
-
财政年份:2005
-
负责人:KUROIWA Shingo
-
依托单位:
海外基金