Person Recognition by Multi-modal Information
Person Recognition by Multi-modal Information
批准号:
09680394
负责人:
KITAMURA Tadashi
金额:
$2.11万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
1997
资助国家:
日本
项目状态:
已结题
起止时间:
1997 至 1998
中文摘要
1. 提出了一种基于语音和面部图像的双峰信息的人脸识别新方法。提出的方法利用隐马尔可夫模型(HMM)对口语单词的唇运动图像序列进行处理。我们研究了强度和位置归一化算法,对双峰数据库Tulips1(12人,4位英文单词)的识别精度达到95%左右。我们还提出了一种新的归一化算法,并表明它比之前提出的算法减少了计算量。我们还将该方法应用于比Tulips1大的双峰数据库M2VTS,该数据库由37人的10位单词组成。在此基础上,研究了基于隐马尔可夫模型的人脸图像归一化和唇部定位跟踪算法。我们只使用唇读信息进行了口语单词识别和说话人识别实验。实验结果表明,使用强度和位置归一化是非常有效的。我们获得了一个单词“0”的说话人识别率为81.0%,对37个人的10个数字的说话人识别率为74.2%。针对基于语音的说话人识别问题,提出了一种利用二阶全通扭曲函数相位特征的频谱参数估计方法。该方法可以改变任意区域内语音频谱的频率分辨率。利用该方法进行了基于区别特征提取(DFE)的说话人识别实验,该方法优化了频谱的扭曲函数,实现了说话人识别。分别用该方法和传统方法进行了说话人识别实验。实验结果表明,该方法比传统方法更有效,2kHz左右的频谱对说话人识别非常重要。
英文摘要
1. We proposed a new technique for person recognition using bimodal information comprising of speech and facial image. The proposed method utilizes a Hidden Markov Model(HMM) for a image sequence of lip movement of a spoken word. We studied intensity and location normalization algorithms and obtained a recognition accuracy of about 95% for a bimodal database Tulips1(12 persons, 4 digit word in English). We also proposed a new normalization algorithm and showed that it reduces the calculation amount less than the one we proposed before.2. We also applied the proposed method to a bimodal database M2VTS bigger than Tulips1, which consists of 10 digit words of 37 persons. Furthermore, some algorithms based on HMM for normalization of facial image and tracking of lip location were studied. We carried out spoken word recognition and speaker identification experiments using only lip reading information. The experimental results have shown that an use of intensity and location normalization is very effective. We obtained a speaker identification rate of 81.0% using one word "0" and a word recognition rate of 74.2% for 10 digits for 37 persons, respectively.3. For speaker identification using speech, we proposed a new spectral parameter estimation method which utilizes a phase characteristics of a second-order all-pass warping function. This method can change the frequency resolution of speech spectrum in an arbitrary region. Using the proposed method we carried out speaker recognition experiments based on a discriminative feature extraction (DFE), which optimizes the warping function of spectrum for speaker recognition. We carried out speaker identification experiments by the proposed method and conventional ones. Experimental results have shown that this method is more effective than conventional methods and spectrum around 2kHz is very important for speaker identification.
期刊论文(76)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Oscar Vanegas etc.: ""HMM-Based Visual Speech Recognition Using Intensity and Location Normalization"" Proceedings of International Conference on Spoken Language Processing (ICSLP98). 289-292 (1998)
Oscar Vanegas等:“使用强度和位置标准化的基于HMM的视觉语音识别”国际口语处理会议论文集(ICSLP98)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Takayoshi Yoshimura etc.: ""State Duration Modeling for HMM-Based Synthesis"" IEICE Tech.Report. SP98-64. 45-50 (1998)
Takayoshi Yoshimura 等:““基于 HMM 的合成的状态持续时间建模””IEICE Tech.Report。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Hiroyuki Inoue etc.: ""Normalization of HMM-Based Speaker Model of Lip Movement for Person Recognition using Bimodal Information"" FSPATJ (Forum of Signal Processing Applications and Technology of Japan) Tech.Report. (1999)
Hiroyuki Inoue 等:“基于 HMM 的嘴唇运动扬声器模型的标准化,用于使用双峰信息进行人物识别” FSPATJ(日本信号处理应用与技术论坛)技术报告。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Jun Hiroi etc.: ""Lip Image Sequence Generation Using HMM"" Proc.of 1999 Spring Meeting of ASJ.2-P-22. 311-312 (1999)
Jun Hiroi 等:“使用 HMM 生成唇部图像序列”Proc.of 1999 Spring Meeting of ASJ.2-P-22。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Tanaka,Vanegas,Tokuda,Kitamura: "Intensity/Location Normalization for Automatic Lipreading" International Conference of Signal Processing. ICSP98. 920-923 (1998)
Tanaka,Vanegas,Tokuda,Kitamura:“自动唇读的强度/位置标准化”信号处理国际会议。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 76 条
Subunit modeling for Japanese sign language recognition based on stochastic model
-
批准号:22500506
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.75万
-
财政年份:2010
-
负责人:KITAMURA Tadashi
-
依托单位:
固有声(eigenvoice)に基づいた音声合成---多様な声質の実現を目指して---
-
批准号:12680380
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.3万
-
财政年份:2000
-
负责人:KITAMURA Tadashi
-
依托单位:
The Studies on the folklore concerning cattle raising in Chugoku mountain area
-
批准号:09610316
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$1.28万
-
财政年份:1997
-
负责人:KITAMURA Tadashi
-
依托单位:
The Formation and Diffusion of the Folk Cultures in Chugoku Mountains
-
批准号:05610250
-
项目类别:Grant-in-Aid for General Scientific Research (C)
-
资助金额:$1.02万
-
财政年份:1993
-
负责人:KITAMURA Tadashi
-
依托单位:
Word Recognition using A Two-Dimensional Mel-cepstrum under Noisy Environments.
-
批准号:63550253
-
项目类别:Grant-in-Aid for General Scientific Research (C)
-
资助金额:$1.34万
-
财政年份:1988
-
负责人:KITAMURA Tadashi
-
依托单位:
On the Value Orientations of Okinawa in the Light of Dynamics of "Monchu" System.
-
批准号:60510151
-
项目类别:Grant-in-Aid for General Scientific Research (C)
-
资助金额:$0.9万
-
财政年份:1985
-
负责人:KITAMURA Tadashi
-
依托单位:
海外基金