Spiral construction of syntactically annotated spoken language corpus

Spiral construction of syntactically annotated spoken language corpus
复制标题

句法注释口语语料库的螺旋构建

DOI:
10.1109/nlpke.2003.1275953
复制
发表时间:
2003
期刊:
International Conference on Natural Language Processing and Knowledge Engineering, 2003. Proceedings. 2003
影响因子:
--
通讯作者:
Y. Inagaki
Y. Inagaki
中科院分区:
--
文献类型:
--
作者:
T. Ohno;S. Matsubara;Nobuo Kawaguchi;Y. Inagaki

文献摘要

被引文献

相似文献

自发言语包括广泛的口语语言现象特征,因此统计方法对口语的鲁棒分析是有效的。随机句法分析需要大规模的语法标注语料库,但其构建需要大量的人力资源。提出了一种高效构建口语语料库的方法,并对语料库进行了依赖分析。这种方法使用现有的口语语料库。采用随机依赖句法分析方法对口语句子进行依赖句法标记,并对结果进行人工校正。标记的语料库以螺旋方式构建,其中校正后的数据用作自动解析其他数据的统计信息。采用这种螺旋方法可以减少解析错误,也可以减少校正成本。用10995个日语话语进行的实验表明,螺旋方法可以有效地构建语料库。
Spontaneous speech includes a broad range of linguistic phenomena characteristic of spoken language, and therefore a statistical approach would be effective for robust parsing of spoken language. Though a large-scale syntactically annotated corpus is required for the stochastic parsing, its construction requires a lot of human resources. We propose a method of efficiently constructing a spoken language corpus for which the dependency analysis is provided. This method uses an existing spoken language corpus. A stochastic dependency parse is employed to tag spoken language sentences with the dependency structures, and the results are corrected manually. The tagged corpus is constructed in a spiral fashion where in the corrected data is utilized as the statistical information for automatic parsing of other data. Taking this spiral approach reduces the parsing errors, also allowing us to reduce the correction cost. An experiment using 10995 Japanese utterances shows the spiral approach to be effective for efficient corpus construction.