Sinica COSPRO and Toolkit — Corpora and Platform of Mandarin Chinese Fluent Speech

Sinica COSPRO and Toolkit — Corpora and Platform of Mandarin Chinese Fluent Speech
复制标题

Sinica COSPRO 及工具包——普通话流利语音语料库及平台

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
Chun
Chun
中科院分区:
--
文献类型:
--
作者:
Chiu;Yun;Chun

文献摘要

被引文献

相似文献

本文报导了中央研究院台北中文连续语音韵律语料库(the Sinica COSPRO)及工具箱的内容及发布情况(http://www.myet.com/COSPRO)。该软件包包括总计约11.99 GB的录制语音语料库(7.7 GB的注释和人工抽查)和开发的分析平台。本文的研究重点是其视角,即如何处理和解释流畅的语音韵律。语料库反映了自上而下的观点,强调在线言语计划和认知限制之间的相互作用关系,以及语音产生过程中的生理和发音限制,仅通过流畅语音韵律中短语内部和短语之间的关联模式来举例说明。由于这些因素之间的相互作用,流利的语音是鲁棒和含糊的语音信号的混合体,不能也不应该简单地将不相关的韵律单位连接成语音串,无论它们是大是小。两个主要特征将我们自上而下的分层多短语框架与大多数其他韵律分析区分开来。一是听者感知到的流利言语的单位和边界;另一个是如何将我们的发现应用到更实际的应用中。我们认为,任何尝试推导或模拟流利的语音韵律必须考虑到上述因素。由于语音数据的收集早在1997年就开始了,一些语料库在O-COCOSDA之前已经被报道过(Tseng et al . 2003)。标注的部分语料库被手工标记为感知的边界和单位,并进行人工抽查。通过COSPRO工具包,我们还与社区分享我们的研究方法。我们的工具包是一个非常友好的基于窗口的平台,它集成了常见的语音分析软件,如Adobe Audition, Praat和speech Viewer到一个平台中。它既接受COSPRO标记的语音文件,也接受用户定义的标签,还可以很容易地合并到现有的数据驱动方法中,以改善韵律输出。该平台包括三个主要功能:(1)进行声学分析;(2)标记连续流畅语音;(3)重新合成语音信号。COSPRO Toolkit最重要的特性是重新合成功能。声学参数可以从COSPRO的韵律单位语音信号中提取,然后单独或集体操作以产生语音输出。基于我们构建的模块化声学模型[1],用户能够操纵F0轮廓,音节持续时间,强度分布和跨短语的边界断裂,无论是在说话者内部还是在说话者之间。因此,Toolkit允许用户改变扬声器的旋律、节奏和语速,或者将上述功能从一个扬声器移植到另一个扬声器以生成不同的韵律输出。我们相信Sinca COSPRO和Toolkit将对语音社区,特别是技术开发非常有用。未来的工作包括进一步研究韵律相关现象,如F0重置和跨短语F0范围的修改模式,以及实现感知边界、单元和跨短语模板的概念,以实现语音识别。
This paper reports the content and release of The Sinica COSPRO (Mandarin Continuous Speech Prosody Corpora) and Toolkit, Academia Sinica, Taipei, Taiwan (http://www.myet.com/COSPRO). The package includes a total of about 11.99 GB recorded speech corpora (7.7 GB annotated and human spot-checked) and analysis platform developed. The focal point of our research is its perspective, namely, how to approach and account for fluent speech prosody. The corpora are reflections of a top-down perspective emphasizing the interacting relationship among on-line speech planning and cognitive limits in addition to physiological as well as articulatory constraints during speech production, exemplified only through the associative patterns within and across phrases in fluent speech prosody. Due to the interactions among these factors, fluent speech is a mixture of both robust and slurred speech signals that can not and should not be viewed simply as concatenation of unrelated prosodic units into speech strings, be they large or small. Two major characteristics distinguish our top-down hierarchical multiple phrase framework from most of other prosody analyses. One is the units and boundaries in fluent speech perceived by listeners; the other is how to implement our findings to more practical applications. We believe that any attempt to derive or simulate fluent speech prosody must account for the above factors. Since speech data collection began as early as 1997, some of the corpora have been reported before at O-COCOSDA (Tseng et al 2003). Annotated part of the corpora was hand labeled for perceived boundaries and units, and human spot-checked. Through the COSPRO Toolkit we also share our research method with the community. Our Toolkit is a very friendly window-based platform that integrates commonly accessible speech analysis software such as Adobe Audition, Praat and Speech Viewer into one platform. It accepts both COSPRO tagged speech files as well as user defined tags and could also be easily incorporated into existing data driven approaches to improve prosody output. The platform consists of three major functions: (1) performing acoustic analysis, (2) labeling continuous fluent speech and (3) re-synthesizing speech signals. The most important feature of COSPRO Toolkit is the re-synthesis function. Acoustic parameters can be extracted from speech signals in prosodic units from COSPRO and subsequently manipulated independently or collectively to generate speech output. Based on a modular acoustic model we constructed [1] , users are able to manipulate F0 contours, syllable durations, intensity distribution and boundary breaks across phrases, both within or across speakers. As a result, the Toolkit allows users to change the melody, rhythm and speaking rate of a speaker, or to port the above features from one speaker to another to generate different prosody output. We believe that Sinca COSPRO and Toolkit will be very useful to the speech community, especially to technology development. Future works include further investigating prosody related phenomena such as F0 reset and modification patterns of F0 range across phrases towards speech synthesis, as well as implementing the concepts of perceived boundaries, units and cross-phrase templates to speech recognition.