Statistical Dependency Parsing in Korean: From Corpus Generation To Automatic Parsing

Statistical Dependency Parsing in Korean: From Corpus Generation To Automatic Parsing
复制标题

韩语统计依存句法分析:从语料库生成到自动句法分析

DOI:
--
复制
发表时间:
2011
期刊:
--
影响因子:
--
通讯作者:
A. Lokasola
A. Lokasola
中科院分区:
--
文献类型:
--
作者:
M. Surbeck;Sally Coxe;A. Lokasola

文献摘要

被引文献

相似文献

本文对朝鲜语依存句法分析作了两方面的贡献。首先,我们从现有的组成树库建立一个韩国依赖树库。对于像韩语这样形态丰富的语言,依存句法分析比成分句法分析具有一些优势。由于没有太多的训练数据可用,我们自动生成依赖树应用头渗透规则和算法的组成树。其次,我们展示了如何从韩语丰富的词法中提取有用的依存分析特征。一旦我们建立了依赖树库,任何统计解析方法都可以应用。具有挑战性的部分是如何从由多个词素组成的标记中提取特征。我们提出了一种方法,选择重要的语素,并只使用这些功能,以避免稀疏。我们的分析方法进行了评估,三个不同的体裁使用黄金标准和自动形态分析。我们还测试了依赖解析的细粒度与粗粒度形态的影响。通过自动形态分析,我们实现了80%+的标记附着评分。据我们所知,这是第一次,韩国的依存关系分析已被评估标记的边缘,这样一个大的各种数据。
This paper gives two contributions to dependency parsing in Korean. First, we build a Korean dependency Treebank from an existing constituent Treebank. For a morphologically rich language like Korean, dependency parsing shows some advantages over constituent parsing. Since there is not much training data available, we automatically generate dependency trees by applying head-percolation rules and heuristics to the constituent trees. Second, we show how to extract useful features for dependency parsing from rich morphology in Korean. Once we build the dependency Treebank, any statistical parsing approach can be applied. The challenging part is how to extract features from tokens consisting of multiple morphemes. We suggest a way of selecting important morphemes and use only these as features to avoid sparsity. Our parsing approach is evaluated on three different genres using both gold-standard and automatic morphological analysis. We also test the impact of fine vs. coarse-grained morphologies on dependency parsing. With automatic morphological analysis, we achieve labeled attachment scores of 80%+. To the best of our knowledge, this is the first time that Korean dependency parsing has been evaluated on labeled edges with such a large variety of data.