Morphological Analysis of a Large Spontaneous Speech Corpus in Japanese

Morphological Analysis of a Large Spontaneous Speech Corpus in Japanese
复制标题

大型日语自发语音语料库的形态分析

DOI:
10.3115/1075096.1075157
复制
发表时间:
2003
影响因子:
15
通讯作者:
H. Isahara
H. Isahara
中科院分区:
化学1区
文献类型:
--
作者:
Kiyotaka Uchimoto;Chikashi Nobata;Atsushi Yamada;S. Sekine;H. Isahara

文献摘要

被引文献

相似文献

本文描述了日语自发语音语料库中词段及其形态信息的两种检测方法,并描述了如何利用这两种方法对大型自发语音语料库进行准确标注。第一种方法用于检测任何类型的词段。第二种方法用于有多个词段及其词类的定义,以及一种类型的词段包含另一种类型的词段。在本文中,我们表明,通过使用半自动分析,我们在检测和标记短单词时达到了99%以上的精度,在检测和标记长单词时达到了97%以上的精度;构成语料库的两类词。我们还表明,使用这两种方法比只使用第一种方法获得更好的精度。
This paper describes two methods for detecting word segments and their morphological information in a Japanese spontaneous speech corpus, and describes how to tag a large spontaneous speech corpus accurately by using the two methods. The first method is used to detect any type of word segments. The second method is used when there are several definitions for word segments and their POS categories, and when one type of word segments includes another type of word segments. In this paper, we show that by using semi-automatic analysis we achieve a precision of better than 99% for detecting and tagging short words and 97% for long words; the two types of words that comprise the corpus. We also show that better accuracy is achieved by using both methods than by using only the first.