Automatic Generation of High Quality CCGbanks for Parser Domain Adaptation

Automatic Generation of High Quality CCGbanks for Parser Domain Adaptation
复制标题

DOI:
10.18653/v1/p19-1013
复制
发表时间:
2019-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Masashi Yoshikawa;Hiroshi Noji;K. Mineshima;D. Bekki
Masashi Yoshikawa;Hiroshi Noji;K. Mineshima;D. Bekki
中科院分区:
其他
文献类型:
--
作者:
Masashi Yoshikawa;Hiroshi Noji;K. Mineshima;D. Bekki

文献摘要

相似文献

我们提出了一种用于组合分类语法(CCG)解析的新域适应方法,基于利用更便宜的依存树资源自动生成 CCG 语料库的思想。我们的解决方案在概念上很简单,并且不依赖于特定的解析器架构,使其适用于当前性能最佳的解析器。我们进行了广泛的解析实验并进行了详细的讨论;在 (1) 生物医学文本和 (2) 问题句子的现有基准数据集之上,我们创建了 (3) 语音对话和 (4) 数学问题的实验数据集。当应用于所提出的方法时,现成的 CCG 解析器显示出显着的性能提升,语音对话从 90.7% 提高到 96.6%,数学问题从 88.5% 提高到 96.8%。
We propose a new domain adaptation method for Combinatory Categorial Grammar (CCG) parsing, based on the idea of automatic generation of CCG corpora exploiting cheaper resources of dependency trees. Our solution is conceptually simple, and not relying on a specific parser architecture, making it applicable to the current best-performing parsers. We conduct extensive parsing experiments with detailed discussion; on top of existing benchmark datasets on (1) biomedical texts and (2) question sentences, we create experimental datasets of (3) speech conversation and (4) math problems. When applied to the proposed method, an off-the-shelf CCG parser shows significant performance gains, improving from 90.7% to 96.6% on speech conversation, and from 88.5% to 96.8% on math problems.