Sinica COSPRO and Toolkit — Corpora and Platform of Mandarin Chinese Fluent Speech
Sinica COSPRO and Toolkit — Corpora and Platform of Mandarin Chinese Fluent Speech
复制标题
Sinica COSPRO 及工具包——普通话流利语音语料库及平台
DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
Chun
中科院分区:
文献类型:
--
作者:
Chiu;Yun;Chun
This paper reports the content and release of The Sinica COSPRO (Mandarin Continuous Speech Prosody Corpora) and Toolkit, Academia Sinica, Taipei, Taiwan (http://www.myet.com/COSPRO). The package includes a total of about 11.99 GB recorded speech corpora (7.7 GB annotated and human spot-checked) and analysis platform developed. The focal point of our research is its perspective, namely, how to approach and account for fluent speech prosody. The corpora are reflections of a top-down perspective emphasizing the interacting relationship among on-line speech planning and cognitive limits in addition to physiological as well as articulatory constraints during speech production, exemplified only through the associative patterns within and across phrases in fluent speech prosody. Due to the interactions among these factors, fluent speech is a mixture of both robust and slurred speech signals that can not and should not be viewed simply as concatenation of unrelated prosodic units into speech strings, be they large or small. Two major characteristics distinguish our top-down hierarchical multiple phrase framework from most of other prosody analyses. One is the units and boundaries in fluent speech perceived by listeners; the other is how to implement our findings to more practical applications. We believe that any attempt to derive or simulate fluent speech prosody must account for the above factors. Since speech data collection began as early as 1997, some of the corpora have been reported before at O-COCOSDA (Tseng et al 2003). Annotated part of the corpora was hand labeled for perceived boundaries and units, and human spot-checked. Through the COSPRO Toolkit we also share our research method with the community. Our Toolkit is a very friendly window-based platform that integrates commonly accessible speech analysis software such as Adobe Audition, Praat and Speech Viewer into one platform. It accepts both COSPRO tagged speech files as well as user defined tags and could also be easily incorporated into existing data driven approaches to improve prosody output. The platform consists of three major functions: (1) performing acoustic analysis, (2) labeling continuous fluent speech and (3) re-synthesizing speech signals. The most important feature of COSPRO Toolkit is the re-synthesis function. Acoustic parameters can be extracted from speech signals in prosodic units from COSPRO and subsequently manipulated independently or collectively to generate speech output. Based on a modular acoustic model we constructed [1] , users are able to manipulate F0 contours, syllable durations, intensity distribution and boundary breaks across phrases, both within or across speakers. As a result, the Toolkit allows users to change the melody, rhythm and speaking rate of a speaker, or to port the above features from one speaker to another to generate different prosody output. We believe that Sinca COSPRO and Toolkit will be very useful to the speech community, especially to technology development. Future works include further investigating prosody related phenomena such as F0 reset and modification patterns of F0 range across phrases towards speech synthesis, as well as implementing the concepts of perceived boundaries, units and cross-phrase templates to speech recognition.