The Role of Context and Uncertainty in Shallow Discourse Parsing

The Role of Context and Uncertainty in Shallow Discourse Parsing
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Katherine Atwell;Remi Choi;Junyi Jessy Li;Malihe Alikhani
Katherine Atwell;Remi Choi;Junyi Jessy Li;Malihe Alikhani
中科院分区:
其他
文献类型:
--
作者:
Katherine Atwell;Remi Choi;Junyi Jessy Li;Malihe Alikhani

文献摘要

相似文献

话语分析已被证明对许多需要复杂推理的NLP任务是有用的。然而,自宾夕法尼亚语篇树库出现十多年以来,预测语篇中的隐含语篇关系仍然具有挑战性。这有几个可能的原因,我们假设模型应该暴露在更多的上下文中,因为它在准确的人类注释中起着重要作用;同时,增加不确定度可以提高模型的精度和定标精度。为了彻底研究这一现象,我们进行了一系列实验,以确定1)语境对人类判断的影响,以及2)用注释者置信度评级量化不确定性对模型准确性和校准的影响(我们使用Brier评分(Brier et al, 1950)进行测量)。我们发现,包括注释器的准确性和置信度可以提高模型的准确性,并且在模型的温度函数中加入置信度可以使模型具有更好的校准置信度。我们还在这些数据集上发现了一些关于人类和模型行为的深刻的定性结果。
Discourse parsing has proven to be useful for a number of NLP tasks that require complex reasoning. However, over a decade since the advent of the Penn Discourse Treebank, predicting implicit discourse relations in text remains challenging. There are several possible reasons for this, and we hypothesize that models should be exposed to more context as it plays an important role in accurate human annotation; meanwhile adding uncertainty measures can improve model accuracy and calibration. To thoroughly investigate this phenomenon, we perform a series of experiments to determine 1) the effects of context on human judgments, and 2) the effect of quantifying uncertainty with annotator confidence ratings on model accuracy and calibration (which we measure using the Brier score (Brier et al, 1950)). We find that including annotator accuracy and confidence improves model accuracy, and incorporating confidence in the model’s temperature function can lead to models with significantly better-calibrated confidence measures. We also find some insightful qualitative results regarding human and model behavior on these datasets.