Integrating Multiple Dependency Corpora for Inducing Wide-coverage Japanese CCG Resources

Integrating Multiple Dependency Corpora for Inducing Wide-coverage Japanese CCG Resources
复制标题

整合多重依存语料库,引入广覆盖的日语CCG资源

DOI:
10.1145/2658997
复制
发表时间:
2013
期刊:
ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP)
影响因子:
--
通讯作者:
Hideki Mima
Hideki Mima
中科院分区:
--
文献类型:
--
作者:
Sumire Uematsu;Takuya Matsuzaki;Hiroki Hanaoka;Yusuke Miyao;Hideki Mima

文献摘要

参考文献

被引文献

相似文献

本文提出了一种面向日语的大范围组合范畴语法资源归纳方法。对于包括英语在内的一些语言,大型注释语料库的可用性和词汇化语法的基于数据的归纳的发展已经实现了深度解析,即,基于词汇化语法的语法分析。然而,日语的深度句法分析还没有得到广泛的研究。这主要是因为大多数日语句法资源是以基于组块的依存结构表示的,而以前的语法归纳方法依赖于树语料库。为了尽可能准确地将语块依存关系中的句法信息转换为短语结构,提出了多个依存关系语料库标注的集成方法。我们的方法首先集成依赖结构和谓词-论元信息,并将它们转换成短语结构树。然后,以与先前提出的方法类似的方式将树转换为CCG推导。转换的质量是经验性的评估方面获得的CCG词典的覆盖率和语法的解析的准确性。虽然在这项研究中使用的转换过程是专门为日本,我们的方法的框架将适用于其他语言的依赖性为基础的分析已被认为是更合适的基于短语结构的分析,由于形态句法功能。
A novel method to induce wide-coverage Combinatory Categorial Grammar (CCG) resources for Japanese is proposed in this article. For some languages including English, the availability of large annotated corpora and the development of data-based induction of lexicalized grammar have enabled deep parsing, i.e., parsing based on lexicalized grammars. However, deep parsing for Japanese has not been widely studied. This is mainly because most Japanese syntactic resources are represented in chunk-based dependency structures, while previous methods for inducing grammars are dependent on tree corpora. To translate syntactic information presented in chunk-based dependencies to phrase structures as accurately as possible, integration of annotation from multiple dependency-based corpora is proposed. Our method first integrates dependency structures and predicate-argument information and converts them into phrase structure trees. The trees are then transformed into CCG derivations in a similar way to previously proposed methods. The quality of the conversion is empirically evaluated in terms of the coverage of the obtained CCG lexicon and the accuracy of the parsing with the grammar. While the transforming process used in this study is specialized for Japanese, the framework of our method would be applicable to other languages for which dependency-based analysis has been regarded as more appropriate than phrase structure-based analysis due to morphosyntactic features.
利用论元位置和类型进行日语谓词论元结构分析
DOI: --
发表时间: 2011
期刊:
影响因子: --
作者:
Yuta Hayashibe;Mamoru Komachi and Yuji Matsumoto
通讯作者: Mamoru Komachi and Yuji Matsumoto
通过基于示例的注释构建的日语助词语料库。
DOI: --
发表时间: 2010
期刊: Proceedings of the Seventh Conference on International Language Resources and Evaluation (LREC'10)
影响因子: --
作者:
Hanaoka H;Mima H;Tsujii J
通讯作者: Tsujii J