The Project for the Corpus of Spontaneous Japanese Spoken by Non-Native Speakers
The Project for the Corpus of Spontaneous Japanese Spoken by Non-Native Speakers
批准号:
17202011
负责人:
TOKI Satoshi
金额:
$30.62万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (A)
财政年份:
2005
资助国家:
日本
项目状态:
已结题
起止时间:
2005 至 2006
中文摘要
非母语日语自发语料库是非母语日语自发语料库,是用于口语研究的大型标注语料库。语料库作为一种新的语言研究数据,已经大规模、高质量地建立起来。然而,大多数语料库针对的是母语人士的演讲,很少针对非母语人士。我们的研究关注非母语者的这类言语,建立了一个大规模的带注释语料库。该语料库可以显示非本族语使用者的语音特征,为中介语和社会语言学的研究提供有益的数据。非母语人士自发日语语料库包含约2000分钟的自发演讲,相当于约36万个单词。所有这些语音材料都使用头戴式近距离交谈麦克风和数据记录,并降采样到16kHz, 16位精度。语音材料是使用双向转录方案,特别为CSJ(自发日语语料库)设计的转录。语音记录有两种不同的转录方式:正字法转录和语音转录。在“正字法”转录中,语音像普通的日语文本一样使用汉字(汉语的象征物)和假名(日语的音节)进行转录,但与普通的日语写作不同,我们的正字法转录对汉字和假名字母的使用有严格的规则。部分语料库被分段标记。这些标签基本上是音位的,但也使用了一些音位标签。语音标签的引入是为了研究语音变异和自发的语音特异性现象。5个语音也用X-JToBI标记语调。在X-JToBI方案中,声调和BI(边界索引)标签都得到了相当大的扩展,以匹配自发语音语调的副语言特征。
英文摘要
The Corpus of Spontaneous Japanese spoken by non-native speakers is a large-scale annotated corpus for spoken language research. Corpus has been focused on as new language research data, which was established with high quality on a large scale. Most of the corpora, however, aim at the speech of native speakers, and few at non-native speakers. Our research, paying attention to such speech of non-native speakers, has established a large-scale annotated corpus. This corpus could show phonetic features in speech of non-native speakers, and furthermore provide research of interlanguage and sociolinguistics with beneficial data.The Corpus of Spontaneous Japanese spoken by non-native speakers contains about 2000 minutes of spontaneous speech that correspond to about 360k words. All these speech material are recorded using head-worn close-talking microphones and DAT, and down-sampled to 16kHz, 16bit accuracy. The speech material is transcribed using a two-way transcription scheme designed especially for CSJ (the Corpus of Spontaneous Japanese).Recorded speech is transcribed in two different ways: orthographic and phonetic transcriptions. In "orthographic" transcription, speech is transcribed using Kanji (Chinese logograph) and Kana (Japanese syllabary) just like ordinary Japanese text, but unlike the ordinary Japanese writing, our orthographic transcription has rigorous rules about the usage of Kanji and Kana letters.Part of the corpus is segment labeled. The labels are basically phonemic, but some phonetic labels are used, too. Phonetic labels are introduced for the study of phonetic variation and spontaneous speech-specific phenomena.5 utterances are also intonation labeled with X-JToBI. In the scheme of X-JToBI both the tone and BI (boundary index) labels were considerably extended to match the paralinguistic features of the spontaneous speech intonation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Research Analysis and Further Advancement of "The Corpus of Spontaneous Japanese Spoken by Non-native Speakers" Project
-
批准号:19320077
-
项目类别:Grant-in-Aid for Scientific Research (B)
-
资助金额:$11.32万
-
财政年份:2007
-
负责人:TOKI Satoshi
-
依托单位:
JAPANESE LANGUAGE AND CULTURE IN MICRONESIA
-
批准号:06041070
-
项目类别:Grant-in-Aid for international Scientific Research
-
资助金额:$13.57万
-
财政年份:1994
-
负责人:TOKI Satoshi
-
依托单位:
A multi-dimentional study in the process and determining factors of acquisition of Japanese as a second language by immigrant workers
-
批准号:06301099
-
项目类别:Grant-in-Aid for Scientific Research (A)
-
资助金额:$3.78万
-
财政年份:1994
-
负责人:TOKI Satoshi
-
依托单位:
Elucidation of a Novel Metabolic Pathway of Morphine and Its Physiological Role
-
批准号:61571078
-
项目类别:Grant-in-Aid for General Scientific Research (C)
-
资助金额:$1.28万
-
财政年份:1986
-
负责人:TOKI Satoshi
-
依托单位:
海外基金