Are All Steps Equally Important? Benchmarking Essentiality Detection in Event Processes

Are All Steps Equally Important? Benchmarking Essentiality Detection in Event Processes
复制标题

DOI:
10.18653/v1/2023.emnlp-main.246
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Hongming Zhang;Yueguan Wang;Yuqian Deng;Haoyu Wang;Muhao Chen;D. Roth
Hongming Zhang;Yueguan Wang;Yuqian Deng;Haoyu Wang;Muhao Chen;D. Roth
中科院分区:
其他
文献类型:
--
作者:
Hongming Zhang;Yueguan Wang;Yuqian Deng;Haoyu Wang;Muhao Chen;D. Roth

文献摘要

相似文献

自然语言以不同的粒度表达事件,粗粒度的事件(目标)可以分解为细粒度的事件序列(步骤)。理解事件过程的一个关键但被忽视的方面是认识到并非所有步骤事件对完成目标都具有同等重要性。在本文中,我们解决这个差距,通过检查在何种程度上,目前的模型理解的步骤事件的重要性有关的目标事件。认知研究表明,这种能力使机器能够模仿人类关于日常任务的先决条件和必要努力的常识推理。我们贡献了一个高质量的语料库(目标,步骤)从社区指南网站WikiHow收集的对,与步骤手动注释的重要性有关的目标由专家。注释者之间的高度一致性表明,人类对事件本质性有着一致的理解。然而,在评估了多个统计和大规模的预训练语言模型后,我们发现现有的方法与人类相比表现不佳。这一意见突出表明,需要进一步探讨这一关键和具有挑战性的任务。数据集和代码可在http://cogcomp.org/page/publication_view/1023上获得。
Natural language expresses events with varying granularities, where coarse-grained events (goals) can be broken down into finer-grained event sequences (steps). A critical yet overlooked aspect of understanding event processes is recognizing that not all step events hold equal importance toward the completion of a goal. In this paper, we address this gap by examining the extent to which current models comprehend the essentiality of step events in relation to a goal event. Cognitive studies suggest that such capability enables machines to emulate human commonsense reasoning about preconditions and necessary efforts of everyday tasks. We contribute a high-quality corpus of (goal, step) pairs gathered from the community guideline website WikiHow, with steps manually annotated for their essentiality concerning the goal by experts. The high inter-annotator agreement demonstrates that humans possess a consistent understanding of event essentiality. However, after evaluating multiple statistical and largescale pre-trained language models, we find that existing approaches considerably underperform compared to humans. This observation highlights the need for further exploration into this critical and challenging task. The dataset and code are available at http://cogcomp.org/page/publication_view/1023.