Universal-Phonetic-Segment-Based Speech Coding and Its Applications to Speech Processing
Universal-Phonetic-Segment-Based Speech Coding and Its Applications to Speech Processing
批准号:
15300026
负责人:
TANAKA Kazuyo
金额:
$10.56万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (B)
财政年份:
2003
资助国家:
日本
项目状态:
已结题
起止时间:
2003 至 2005
中文摘要
在这个项目中,我们提出了一个新的语音处理框架,其中所有的声学语音样本一次编码成通用语音段(UPS)序列和口语文档处理(SDP)系统,如识别,检索,索引,构建在这个UPS域。采用该框架,SDP系统从原始声学相关物或环境分离。这使得可以实现这样的灵活性,即可以通过计算UPS序列之间的距离来处理简化型处理,也可以构建分布式处理方案。通过该项目,我们在该框架上开发了以下组件技术:1)采用原始的细亚语音段(SPS)集作为UPS集,实现了多语种语音的高性能识别和简单处理,2)有效的基于动态规划(DP)的序列匹配算法,称为移位CDP和中继CDP。的处理框架,SPS集,和DP为基础的算法的有效性进行评估,通过构建语音识别和开放词汇口语文档检索(SDR)系统。实验结果表明,所提出的SDP系统的性能评价优于那些基于传统方法的上级。最后,我们构建了一个真实的实时开放词汇表SDR系统进行演示,该系统可以通过用户的语音检索广播视频。
英文摘要
In this project, we present a novel speech processing framework, where all of the acoustic speech samples are once encoded into universal phonetic segment (UPS) sequences and spoken document processing (SDP) systems, such as recognition, retrieval, indexing, are constructed on this UPS domain. Adopting this framework, the SDP systems are separated from the original acoustic correlates or environments. This makes it possible to realize such flexibility that recognition-type processing can be handled by just calculating distances between UPS sequences, and also can be constructed on distributed processing schemes.Through this project, we have developed the following component techniques on this framework : 1)an original fine sub-phonetic segment (SPS) set as the UPS set, which brought high performance recognition and easy processing of multilingual speech, 2)effective DP(dynamic programming)-based sequence matching algorithms, called Shift CDP and Relay CDP. Effectiveness of the processing framework, the SPS set, and DP-based algorithms are evaluated by constructing speech recognition and open vocabulary spoken document retrieval (SDR) systems. Experimental results showed that the proposed SDP systems are superior to those based on conventional methods in performance evaluation. We have finally constructed a real time open vocabulary SDR system for demonstration, in which the system can retrieve broadcast video by user's speech.
期刊论文(87)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
音素片のカーネル主成分分析を用いたトピックセグメンテーション
使用音素片段的核主成分分析进行主题分割
DOI:
--
发表时间:
2005
期刊:
人工知能学会 1E2-03
影响因子:
--
作者:
[佐土原健, 児島宏明, 李時旭]
通讯作者:
李時旭
Similar section extraction for analyzing stream data structure
用于分析流数据结构的相似部分提取
DOI:
--
发表时间:
2004
期刊:
Proc.of 5th European Conference on Machine Learning (ECML2004) 1
影响因子:
--
作者:
[Itoh, Y., Tanaka, K., Lee, S.W.]
通讯作者:
S.W.
HMM-Based Feature Compensation Method : An Evaluation Using the AURORA2
基于 HMM 的特征补偿方法:使用 AURORA2 进行评估
DOI:
--
发表时间:
2004
期刊:
Proc.of International Conference on Spoken Language Processing (ICSLP2004) 1
影响因子:
--
作者:
[Sasou, A., Asano, F., Tanaka, K., Nakamura, S.]
通讯作者:
S.
Mixed-Lingual Spoken Word Recognition by Using VQ Codebook Sequnces of Variable Length Segments
使用变长片段VQ码本序列的混合语言口语单词识别
DOI:
--
发表时间:
2003
期刊:
Proc. of the European Conference on Speech Communication and Technology 4
影响因子:
--
作者:
[Kojima, H.]
通讯作者:
H.
A fast matching algorithm called shift continuous DP between arbitrary parts of two time sequence data sets,
两个时序数据集任意部分之间的称为移位连续DP的快速匹配算法,
DOI:
--
发表时间:
2003
期刊:
IEICE Trans.Information and Systems (Japanese Ed.) Vol.J89-D, No.3
影响因子:
--
作者:
[山中俊介, 栗林祐宇真, 山子佳久, 宮内宏之, 松島恭治, Shigeki Okawa, Kazuyo Narita, 坂本雄児, Yoshiaki Itoh]
通讯作者:
Yoshiaki Itoh
共 74 条
Development of Continuous Voice Morphing Using Separated Vocal TractArea Functions, Glottal Source Waves, and Prosodic Features
-
批准号:22500145
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.66万
-
财政年份:2010
-
负责人:TANAKA Kazuyo
-
依托单位:
Realtime and multiple degree-of-freedom electric hand system based on electromyogram signals
-
批准号:19500377
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.91万
-
财政年份:2007
-
负责人:TANAKA Kazuyo
-
依托单位:
海外基金