CPTAM: Constituency Parse Tree Aggregation Method

CPTAM: Constituency Parse Tree Aggregation Method
复制标题

CPTAM:选区解析树聚合方法

DOI:
10.1137/1.9781611977172.71
复制
发表时间:
2022
期刊:
Proceedings of the SIAM International Conference on Data Mining
影响因子:
--
通讯作者:
Li, Qi
Li, Qi
中科院分区:
--
文献类型:
--
作者:
Kulkarni, Adithya;Sabetpour, Nasim;Markin, Alexey;Eulenstein, Oliver;Li, Qi

文献摘要

参考文献

相似文献

各种自然语言处理任务根据短语结构语法使用成分分析来理解句子的句法结构。人们提出了许多最先进的选区句法分析器,但对于相同的句子,它们可能会提供不同的结果,特别是对于其训练领域以外的语料库。本文采用真值发现的思想,在缺乏基本真值的情况下,通过估计不同句法分析器的可信度来聚合不同句法分析器的选区句法分析树。我们的目标是始终如一地获得高质量的聚集选区分析树。我们将选区分析树聚合问题分为结构聚合和成分标签聚合两个步骤进行描述。具体地说,我们提出了通过最小化Robinson-Foulds(RF)距离的加权和来发现树结构的第一个真理发现解决方案,RF距离是两棵树之间的经典对称距离度量。在不同语言和领域的基准数据集上进行了广泛的实验。实验结果表明,我们的方法CPTAM的性能优于最新的聚合基线。我们还证明了在缺乏基本事实的情况下,CPTAM估计的权重可以充分地评估选区分析器。
Diverse Natural Language Processing tasks employ constituency parsing to understand the syntactic structure of a sentence according to a phrase structure grammar. Many state-of-the-art constituency parsers are proposed, but they may provide different results for the same sentences, especially for corpora outside their training domains. This paper adopts the truth discovery idea to aggregate constituency parse trees from different parsers by estimating their reliability in the absence of ground truth. Our goal is to consistently obtain high-quality aggregated constituency parse trees. We formulate the constituency parse tree aggregation problem in two steps, structure aggregation and constituent label aggregation. Specifically, we propose the first truth discovery solution for tree structures by minimizing the weighted sum of Robinson-Foulds (RF) distances, a classic symmetric distance metric between two trees. Extensive experiments are conducted on benchmark datasets in different languages and domains. The experimental results show that our method, CPTAM, outperforms the state-of-the-art aggregation baselines. We also demonstrate that the weights estimated by CPTAM can adequately evaluate constituency parsers in the absence of ground truth.
DOI: 10.18653/v1/w18-2501
发表时间: 2018-03
期刊: ArXiv
影响因子: --
作者:
Matt Gardner;Joel Grus;Mark Neumann;Oyvind Tafjord;Pradeep Dasigi;Nelson F. Liu;Matthew E. Peters
通讯作者: Matt Gardner;Joel Grus;Mark Neumann;Oyvind Tafjord;Pradeep Dasigi;Nelson F. Liu;Matthew E. Peters
OptSLA:一种基于优化的顺序标签聚合方法
DOI: --
发表时间: 2020
期刊: Findings of the Association for Computational Linguistics: EMNLP 2020
影响因子: --
作者:
Sabetpour, Nasim;Kulkarni, Adithya;Li, Qi
通讯作者: Li, Qi
DOI: 10.18653/v1/d16-1180
发表时间: 2016-09
期刊: Sensors (Basel, Switzerland)
影响因子: --
作者:
A. Kuncoro;Miguel Ballesteros;Lingpeng Kong;Chris Dyer;Noah A. Smith
通讯作者: A. Kuncoro;Miguel Ballesteros;Lingpeng Kong;Chris Dyer;Noah A. Smith
DOI: 10.1136/ebmh.11.4.102
发表时间: 2008-10
期刊: Evidence Based Mental Health
影响因子: --
作者:
P. Cochat;L. Vaucoret;J. Sarles
通讯作者: P. Cochat;L. Vaucoret;J. Sarles
用于依存句法分析的反向修正和线性树组合
DOI: 10.3115/1620853.1620925
发表时间: 2009
期刊: Proceedings of the 2005, American Control Conference, 2005.
影响因子: --
作者:
Giuseppe Attardi;F. Dell’Orletta
通讯作者: F. Dell’Orletta