Improving Neural RST Parsing Model with Silver Agreement Subtrees

Improving Neural RST Parsing Model with Silver Agreement Subtrees
复制标题

DOI:
10.18653/v1/2021.naacl-main.127
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Naoki Kobayashi;T. Hirao;Hidetaka Kamigaito;M. Okumura;M. Nagata
Naoki Kobayashi;T. Hirao;Hidetaka Kamigaito;M. Okumura;M. Nagata
中科院分区:
其他
文献类型:
--
作者:
Naoki Kobayashi;T. Hirao;Hidetaka Kamigaito;M. Okumura;M. Nagata

文献摘要

被引文献

相似文献

先前的大多数修辞结构理论(RST)解析方法基于神经网络等监督学习,这需要有足够规模和质量的标注语料库。然而,英语中用于RST解析的基准语料库——修辞结构篇章树库(RST - DT)由于RST树的标注成本高昂而规模较小。缺乏大规模标注训练数据导致性能不佳,尤其是在关系标注方面。因此,我们提出一种通过利用“银数据”(即自动标注的数据)来改进神经RST解析模型的方法。我们使用最先进的RST解析器从未标注的语料库中创建大规模的银数据。为了获得高质量的银数据,我们从使用RST解析器构建的文档的RST树中提取一致子树。然后,我们使用获得的银数据对神经RST解析器进行预训练,并在RST - DT上对其进行微调。实验结果表明,我们的方法在核心性和关系方面分别取得了最佳的微F1分数,分别为75.0和63.2。此外,与之前最先进的解析器相比,我们在关系分数上获得了显著的提高,提高了3.0分。
Most of the previous Rhetorical Structure Theory (RST) parsing methods are based on supervised learning such as neural networks, that require an annotated corpus of sufficient size and quality. However, the RST Discourse Treebank (RST-DT), the benchmark corpus for RST parsing in English, is small due to the costly annotation of RST trees. The lack of large annotated training data causes poor performance especially in relation labeling. Therefore, we propose a method for improving neural RST parsing models by exploiting silver data, i.e., automatically annotated data. We create large-scale silver data from an unlabeled corpus by using a state-of-the-art RST parser. To obtain high-quality silver data, we extract agreement subtrees from RST trees for documents built using the RST parsers. We then pre-train a neural RST parser with the obtained silver data and fine-tune it on the RST-DT. Experimental results show that our method achieved the best micro-F1 scores for Nuclearity and Relation at 75.0 and 63.2, respectively. Furthermore, we obtained a remarkable gain in the Relation score, 3.0 points, against the previous state-of-the-art parser.